Scope¶
Scope decides which datasets the rule is offered.
It is a selection, never a verdict. A dataset outside a rule's scope is not judged by it — which is a different outcome from being judged and found conforming, and different again from being skipped for a missing requirement.
Scope:
Classes:
Include: ["FINDINGS"]
⭐ Scope by the concept, not by a list of datasets¶
Prefer the most general axis that expresses the requirement. If the rule is about every findings domain, write the class:
# ✅ says what the rule is about
Scope:
Classes:
Include: ["FINDINGS"]
# ⛔ says which datasets happened to exist when it was written
Scope:
Domains:
Include: ["LB", "VS", "EG", "PC", "PP", "MB", "IE", "QS", "SC", "FT"]
The two may select the same datasets today. They behave differently the moment the standard changes.
What happens when the standard adds a dataset¶
A new findings domain is published. The class-scoped rule is offered it immediately and keeps enforcing the requirement. The domain-list rule is not — and nothing reports that. No error, no warning, no gap in a coverage report. The rule keeps passing, on a population that quietly stopped being the population it was written for.
A stale scope fails silently, and silence is indistinguishable from conformance. This is the argument for the general axis: it is not about brevity, it is about which way the rule fails when the world moves underneath it.
When a narrow scope is the right answer¶
Scope narrowly when the requirement is about that dataset's own semantics, not about a property its class shares:
TS— the rule is about a named trial-summary parameter. No other dataset hasTSPARMCD.DM— the rule is aboutRFSTDTCas the study reference start. Another special-purpose dataset acquiring a similar column would not make the rule apply to it.RELREC,SUPP--— the rule is about the relationship mechanism itself.
The test is a question about the future: if the standard added another dataset of this class tomorrow, should this rule judge it? If yes, scope by class or structure. If no, name the dataset — and the narrow scope is then a statement, not an accident.
The axes come in two families¶
Classes and Domains are SDTM / SEND terminology. Data_Structures and Subclasses are
ADaM. They are not alternatives to each other; they describe different standards.
| standard | broad axis | narrow axis |
|---|---|---|
| SDTM / SEND | Classes |
Domains |
| ADaM | Data_Structures |
Subclasses |
Two further axes cut across both:
| axis | selects by |
|---|---|
Datasets |
the dataset name, exactly |
Use_Case |
an externally supplied use case |
Every axis you give must be satisfied. They narrow; they do not offer alternatives. Mixing
families — a rule with both Classes and Data_Structures — asks for a dataset that is
simultaneously SDTM and ADaM, and selects nothing.
Give one axis. A second is right only when it genuinely narrows something the first cannot express.
Include and Exclude¶
Use one facet or the other — never both on the same axis. Each already implies the other:
- an
Includeexcludes everything it does not name; - an
Excludeincludes everything it does not name.
# ⛔ pointless — AE was never in scope to begin with
Domains:
Include: ["LB", "VS"]
Exclude: ["AE"]
# ⛔ also pointless — the Exclude alone says exactly this
Domains:
Include: ["ALL"]
Exclude: ["RELREC"]
# ✅ every domain except one
Domains:
Exclude: ["RELREC"]
An axis with no Include matches everything the Exclude does not remove, so "all but these"
is written with an Exclude on its own.
⛔
ALLis not allowed in anExclude, and on the closed ADaM axes the loader rejects it outright: it can never match a detected token, so the exclusion would be a silent no-op.
Include: ["ALL"]¶
Selecting every dataset is the default — a rule with no Scope at all is offered everything.
Include: ["ALL"] is therefore not required.
Write it anyway when that is the intent. It says "this rule is deliberately about every
dataset", where an absent Scope says only "nobody wrote one". The two look identical to the
engine and completely different to a reviewer.
⚠ That is the one place ALL belongs: on its own, as the whole scope. Not beside an Exclude.
Domains — the SDTM / SEND narrow axis¶
Scope:
Domains:
Include: ["AE"]
What an entry may be¶
| entry | matches |
|---|---|
AE |
the AE domain, literally |
ALL |
every domain |
NONE |
no domain — an explicit empty selection |
SUPP-- |
a glob: every SUPP domain (SUPPAE, SUPPDM, …) |
/^AD.*$/ |
a regular expression, between slashes |
-- |
the strict domain-prefix token |
Split domains¶
A large domain may be delivered as several files — LB1, LB2 — that together are one LB.
Scope.Domains is split-aware: Include: ["LB"] selects LB1 and LB2 through the
data-derived unsplit name, not just a file literally called LB.
include_split_datasets makes that explicit:
true— the rule applies only where the domain is actually split. A rule comparing split members needs this; without it the rule also runs on unsplit domains, where it has nothing to compare.false— the rule does not apply to split members.- absent — both are in scope, which is what almost every rule wants.
Classes — the SDTM / SEND broad axis¶
Scope:
Classes:
Include: ["FINDINGS"]
The vocabulary in use:
FINDINGS · FINDINGS ABOUT · INTERVENTIONS · EVENTS · SPECIAL PURPOSE · RELATIONSHIP ·
TRIAL DESIGN
⚠
FINDINGSincludesFINDINGS ABOUT. A rule scoped toFINDINGSis also offeredFINDINGS ABOUTdatasets, deliberately — the latter is a specialisation of the former. If your rule must not see them, exclude them explicitly; the include alone will not do it.
Data_Structures and Subclasses — the ADaM axes¶
Scope:
Data_Structures:
Include: ["BASIC DATA STRUCTURE"]
Data_Structures selects on the dataset's detected ADaM structure. The vocabulary is
closed — exactly these eight tokens, plus the ALL sentinel in Include:
SUBJECT LEVEL ANALYSIS DATASET · BASIC DATA STRUCTURE · OCCURRENCE DATA STRUCTURE ·
MEDICAL DEVICE BASIC DATA STRUCTURE · MEDICAL DEVICE OCCURRENCE DATA STRUCTURE ·
DEVICE LEVEL ANALYSIS DATASET · REFERENCE DATA STRUCTURE · ADAM OTHER
Subclasses selects on the detected subclass, from the Define-XML 2.1 vocabulary — five tokens,
plus ALL:
ADVERSE EVENT · MEDICAL DEVICE TIME-TO-EVENT · NON-COMPARTMENTAL ANALYSIS ·
POPULATION PHARMACOKINETIC ANALYSIS · TIME-TO-EVENT
Write the token exactly — short forms are not accepted¶
⛔
OCCDS,BDSandADSLare not tokens. They are the names the engine's source code uses for these constants; the authored value is the full string. WriteOCCURRENCE DATA STRUCTURE, notOCCDS.
The match is by exact string, so the spelling matters in full: not an abbreviation, not
Basic Data Structure in mixed case, not BASIC_DATA_STRUCTURE.
Getting it wrong fails at load, and that is the point. These vocabularies are closed
precisely so a wrong token cannot slip through — an unrecognised entry can never match a detected
value, so it would silently disable the rule (in Include) or do nothing at all (in Exclude).
The loader rejects it by name instead.
⚠ The canonical key spelling is Data_Structures. The upstream Data Structures, with a space,
is accepted as an alias — prefer the canonical one.
A specialisation matches its base¶
The structure tokens form an is-a hierarchy, and so do the subclasses. A dataset detected as
MEDICAL DEVICE BASIC DATA STRUCTURE also matches a scope naming BASIC DATA STRUCTURE.
So scoping to a base token covers its specialisations — the same principle as
FINDINGS covering FINDINGS ABOUT, and the same reason
to prefer it: a specialisation added to the standard later is picked up by a base-scoped rule and
missed by one that enumerates the specialised tokens.
The asymmetry that catches people¶
Both are detected, and a dataset may have no detectable subclass at all — the normal case for
a plain BDS or ADSL dataset. Include and Exclude treat that case oppositely:
Includerequires a positive detection. A dataset whose subclass cannot be detected is not selected, so anIncludenarrows harder than it looks.Excluderejects only on a positive match. A dataset whose subclass cannot be detected passes an exclude-only scope.
If your rule should judge every dataset except one subclass, use
Exclude. Writing the complement as anIncludelist silently drops every dataset whose subclass is undetectable.
Datasets — the pure dataset name¶
Scope:
Datasets:
Include: ["ADSL"]
This is not Domains under another name. It matches the dataset's name, exactly, with no
domain interpretation and no split re-test: Datasets: ["LB"] selects the dataset called
LB and nothing else, where Domains: ["LB"] would also select LB1 and LB2.
It is the axis that matters for ADaM. ADaM datasets are not domains — ADSL, ADAE, ADTTE
are names. A rule about one specific analysis dataset names it here; Domains has no meaning for
it.
The entry vocabulary is the same as Domains — literals, ALL, NONE, globs, /regex/.
⛔
include_split_datasetsdoes not exist on this axis, and writing it is caught by the loader. It is a statement about domain families; on a name axis it has no meaning.⚠ Glob,
/regex/andNONEentries here are a coreJ extension — legal, and not portable to the upstream CORE schema, whoseDatasetsaxis is a name enum plus a name pattern.
Use_Case¶
Scope:
Use_Case: "INDH"
A use case is supplied externally, by whoever runs the validation — it is not a property of the data. It lets one corpus serve several regulatory contexts from the same rules.
The matching rule has three arms:
| no use case supplied to the run | every rule is in scope; the axis does nothing |
| a use case is supplied, and the rule declares none | the rule is in scope |
| a use case is supplied, and the rule declares some | in scope only if one of them is exactly the supplied one |
So declaring a Use_Case narrows a rule to that context, and leaving it absent keeps the rule
universal. Matching is case-insensitive.
Several use cases¶
A rule may name more than one. They are written as a comma-separated string:
Use_Case: "INDH, PROD"
⚠ Not a YAML array. The field binds as a single string and is split on commas at match time;
Use_Case: ["INDH", "PROD"]is not the shape the engine reads. Whitespace around the commas is trimmed.
Things that are not scope¶
⛔
Scope.Variablesdoes not exist. A rule that writes one is rejected on load. Variable presence is a requirement, not a scope — see Requirements. The rejection exists precisely so the requirement cannot silently not-exist.⛔ Scope is not a row filter. It decides which datasets the rule sees, never which rows. Restricting rows is the check's job, or a precondition's.
Any key under Scope that is not one of the six axes is recorded by the loader and reported, so a
misspelled axis fails loudly rather than binding to nothing and leaving the rule unscoped.
Next: Requirements — what must exist before the rule runs.