Every check, in the open.

A validation engine for standardised clinical data: rule sets for CDISC (SDTM, ADaM, SEND), FDA and PMDA, checks written in a language you can read, understand and extend. That makes it easy to add custom checks, for example against a company-extended implementation guide.

Open source · AGPL-3.0 CDISC · FDA · PMDA Inside Cumba Data Browser
CoreJ - ein Herz, dessen rechte Haelfte in ein j auslaeuft

What CoreJ is

A fast engine

A runtime engine that checks clinical data, written in Java and open source. Rules within each dataset run in parallel on as many worker threads as you choose, and every check is compiled once when it is loaded, then run a whole column at a time rather than row by row.

Complete rule sets

CDISC: SDTMIG 3.2–3.4, ADaMIG 1.0–1.3, SENDIG 3.0–3.1.1, SEND-DART 1.1–1.2, SEND-GENETOX 1.0. FDA: SDTMIG 3.1.2–3.4, SENDIG 3.0–3.1.1, SEND-AR 1.0, SEND-DART 1.1. PMDA: SDTMIG 3.1.2–3.4, ADaMIG 1.0–1.3. Plus Define-XML 2.0 and 2.1 for CDISC and PMDA. Ready to apply: only published rules ship.

A language for checks

The rule sets that ship with CoreJ and the checks you add yourself are written in the same language: plain expressions over the dataset's variables, with functions for the four kinds of check.

  • normal checks — values, text, patterns
  • clinical dates and durations
  • keys and relationships
  • across rows and datasets

Your own rules alongside

Company-specific checks in that same language, running in the same pass as the standard ones.

Grouped findings

Where the rule allows it, a finding is one entry per rule and dataset with its count — not one line per row. The affected rows, variables and values stay available underneath.

Three comparisons

Conformance is checked in three directions — structure, labels, lengths, codelists. Each one optional:

  • data against your Define-XML
  • data against the CDISC definitions
  • your Define-XML against the CDISC definitions

Define-XML 2.0 and 2.1.

Tested rule by rule

Nearly nine in ten rules ship with their own test data (cases that must be flagged and cases that must pass), and every build runs them. Each rule carries the id its authority, CDISC, FDA or PMDA, publishes it under.

One language for every check

The rule sets that ship with CoreJ and the checks you add yourself are written in the same language.

Example check
Core: { Id: "P300-AE0001" }
Description: "AESTDTC is earlier than RFICDTC in DM."
Scope: { Domains: { Include: ["AE"] } }
Match_Datasets: [ { Name: "DM", Keys: ["USUBJID"] } ]
Check:
  expression: >-
    not empty(AESTDTC) and date(AESTDTC) < date(DM.RFICDTC)
Outcome:
  Message: "AESTDTC is before RFICDTC."
  Output_Variables: ["AESTDTC", "DM.RFICDTC"]

Add checks of your own for what your organisation requires — same language, same run. Write your own: the rule guide →

Example finding
Core Id:  P300-AE0001
Dataset:  ae.dsjc      Row: 412
Variable: AESTDTC
Value:    "2021-03-04"
Message:  AESTDTC is before RFICDTC.

The finding carries the check's id and its message, so the whole rule can be read from it.

How can I use CoreJ?

Three ways, from the one that needs no setup to the one that puts the engine inside your own code.

1

In Cumba Data Browser

One click runs the checks on the data you have open, and the findings are highlighted at the rows and values they belong to, or listed in tabular form for processing and documentation. No setup beyond the browser itself.

cumba.net →

2

As a library, in your own tool

Call the engine from your own Java code: an in-house review tool, a data pipeline, your own front end. You hand it the datasets and the rule sets to run, and get the findings back. Cumba Data Browser uses this path.

The engine is on Maven Central as net.cumba:cumba-oss-corej-core, with the report writers cumba-oss-corej-report-json and cumba-oss-corej-report-xlsx beside it. You describe the run — rule packages, datasets, Define-XML, controlled terminology — start the validation, and get the findings back as an object. Java 25; give large studies enough memory, for example -Xmx8g.

Maven Central →   Source on GitHub →

3

On the command line

Run a study from a script, a pipeline step or a nightly job. The data stays on the machine that runs the check.

Download the release, unzip it and run it — the rule sets come bundled, and Java 25 is all it needs: ./run.sh -rp cdisc-sdtmig-3-4 -d ./datasets -o report.json. The findings come out as JSON, as Excel (-of xlsx), or both.

Releases on GitHub →

CoreJ was developed for Cumba Data Browser, and we open-sourced the whole engine under AGPL-3.0. Anyone may use it, in their own company and in their own study. Building it into a product you offer to others means opening that product's source in turn.

Tell us how it goes

A rule that fired when it should not have. One that stayed quiet when it should have spoken up. A check you needed and did not find. Any of those is worth a message, and none of them is too small.

Found a bug? Send the rule, the input that triggered it and what you expected instead — that is what we need to reproduce it.

Need the documentation? Each repository's README on GitHub, and the rule guide for writing checks of your own. A manual is on its way.

Want to write to us? Pick what fits, the email opens ready to go:

Want to stay posted? Releases go out in the P300 newsletter, one for all our tools.