Technical Note 002 · 12 August 2026 · v1.0

286 Pages of Candor, and Still No Assembled u_D

SHIIVER — the best-documented cryogenic tank test I have read — and what a modeler still cannot reconstruct from it.

doi:10.5281/zenodo.21895647 · published in Zenodo, CC BY 4.0

What this note is: a close reading of the SHIIVER final report and its supporting documents, applied against the six referent-quality criteria of this series — first published in Technical Note 001, canonical wording frozen in Technical Note 000 (doi:10.5281/zenodo.21895568). Every factual claim about a specific document is traceable to a passage in that document, listed in Section A of the sources. (“Audit” here means documentary audit, as defined in Note 000.)

What it is not: a review of whether SHIIVER met its own test objectives, which is NASA’s business and not mine — and not a claim that this referent is flawless. The argument of this note is that its flaws are documented, which is a different and better thing.


I. The test

Between August 2019 and January 2020, NASA tested a 4-meter-diameter stainless steel cryogenic tank — the Structural Heat Intercept, Insulation, and Vibration Evaluation Rig, SHIIVER — at Glenn Research Center’s Plum Brook Station in Sandusky, Ohio, since renamed the Neil A. Armstrong Test Facility (the final report states both names, p. 13). Thermal-vacuum testing used the In-Space Propulsion Facility; acoustic exposure, the Reverberant Acoustic Test Facility. The rig was built as a testbed for scaling cryogenic fluid management technologies to large upper stages: a tank, structural skirts supporting it aft, and an aluminum forward skirt with vapor-cooling channels bolted on.

It was tested with liquid hydrogen, and with liquid nitrogen “as a substitute fluid for liquid oxygen and liquid methane.” The campaign ran in four stages: a baseline thermal-vacuum test before multilayer insulation was installed, a thermal-vacuum test with MLI, a reverberant acoustic test at 147 dB overall sound pressure level, and a final thermal-vacuum test to verify the acoustic exposure had done no damage. Each thermal-vacuum test ran continuously from roughly 90 percent full down to 25 percent, with chamber walls at ambient temperature and vacuum in the 10⁻⁶ torr range.

The headline results, from the report’s own summary: vapor cooling of the forward skirt reduced the heat load to the tank by approximately 10 percent, but boiloff by less than 3 percent at 50 percent fill without MLI on the domes — and left boiloff “essentially unchanged” at that fill level once dome MLI was installed. The MLI itself reduced heat load by approximately 40 percent at all fill levels, and boiloff by approximately 25 percent at 90 percent full and 45 percent below 65 percent full. The Radio Frequency Mass Gauge tracked fluid mass across the full fill range.

Those numbers are interesting. This note is about something else: the documentation posture of the test — because in Note 001 I traced what did and did not travel from a 1992 experiment into the validation papers examined there, and SHIIVER is the natural counter-case. It is the most thoroughly documented large-scale cryogenic tank test I have read. The question is what “most thoroughly documented” buys a modeler who wants to validate against it — and what it still does not.

The final report, NASA/TP-20205008233, runs 286 pages and was written by an eight-author team spanning NASA Glenn, Case Western Reserve University, the Universities Space Research Association, and the Jacobs Space Exploration Group.


II. What good looks like

Six practices in the SHIIVER documentation deserve to be named, because together they form a checklist that most of the referents audited in this registry — including ones the community has validated against for decades — do not meet simultaneously.

1. The measurement-performance requirements live in a standalone test plan. (A requirement set in advance is not a demonstrated achieved uncertainty — a requirement, a vendor accuracy specification, a calibration result and a propagated result-level uncertainty are four different objects; this series’ criterion-3 scale keeps them separate.) A 53-page test plan (Revision I, released 9 January 2020 — its own revision history records updates “based on actuals” from the 2019 test phases, so whether each specific figure predates the measurement it governed would require the pre-test revision, not obtained) was published separately from the results. It specifies, sensor by sensor, what measurement performance the test needed: vent-line temperature sensors better than ±0.1 K, because — as the plan puts it — “Fluid enthalpy are very temperature dependent”; tank pressure better than ±0.02 psia, because of “the sensitivity of fluid state to tank pressure”; vapor-cooling exit pressure better than 0.01 psia; boiloff flow meters — the quantity from which heat load is derived — better than “1% of the measurement,” with the plan noting that flow meter uncertainty is “extremely important.” Reporting achieved accuracy afterwards is good practice. Committing to required accuracy beforehand is a different discipline, and the right one.

And the plan does something more, which deserves its own sentence: it declares, prospectively, that the data would serve modelers. “The data from SHIIVER will be used multiple ways, from basic analysis to verification of complex computational models” (test plan, p. 13), and “The temperature rake is used for anchoring computational models to the test data” (p. 23). SHIIVER was explicitly designed to provide data for computational-model comparison and anchoring; calling that a validation referent is this series’ V&V terminology, not the plan’s — the interpretive step is ours and is stated here. Nothing else in this note reinterprets the rig after the fact — which makes the question of what a modeler can reconstruct from its record the exact question its own authors invited.

2. The analysis property basis is explicit. Appendix I of the final report tabulates thermophysical properties for parahydrogen and nitrogen, with references. A modeler does not have to guess which hydrogen the analysis assumed. Note 001 documented that the K-Site report of 1992 contains no statement of ortho/para composition at all; here the assumption is printed, sourced, and checkable. (The spin composition of the fluid actually loaded is a separate matter, and remains unstated — the scorecard below draws that distinction.)

3. The measurements are redundant and dissimilar — and the report shows their disagreements. Fill level and boiloff flow were measured redundantly, by dissimilar principles where available: a capacitance probe, a rake of silicon diodes, four flowmeters — FM1 on the vapor-cooling line; FM2, FM3 and FM4 in the vent line, plumbed in series so that “the same flow passes through multiple flowmeters” — and the Radio Frequency Mass Gauge. The report then does the honest thing: it plots them against each other and narrates the discrepancies (pp. 130–131). FM2 “reads a higher flow rate than FM3 for all tests.” FM1’s performance “was erratic throughout this test sequence” — reading significantly high during one test, then returning to consistency after a pressure-rise test, an anomaly the report flags rather than smooths. Flow-rate deviations between methods are “typically less than 1 g/s” with “several spikes where the deviation is as large as 3 g/s.” Fill-level deviations between the four methods reach “as high as 4 percent.” The documentation also identifies low-flow registration limits in the flowmeters; the exact FM3 cutoff appears as ~2 g/s in one passage and 0.75 g/s in another, and is left unresolved here (the contradiction is recorded in the verification log).

4. The assumptions are stated as assumptions. When the boiloff flow analysis anchors on flowmeter FM3, the report writes, in parentheses, “(it is assumed that FM3 is accurate)” (p. 130). Seven words. They cost the authors nothing and they tell a future modeler exactly where the analysis chain is anchored and what would invalidate it.

5. A 50 percent uncertainty was published rather than hidden. The heat flux sensors mounted on the tank were calibrated — selected units, to ASTM C-1130 and C-1774, before installation — and the team found the published vendor sensitivities usable at liquid hydrogen temperatures but with “uncertainties on the order of 50%” (heat flux paper, p. 3 of 8). The physics is explained: the sensors are thermopiles, the differential voltage they generate at 20 K is small, and at low heat flux the uncertainty grows. The companion paper publishes the number anyway and argues the values and trends still “match well with other calculation methods”. A 50 percent uncertainty stated plainly is worth more to a downstream user than a clean-looking number with no uncertainty at all — and it is exactly the information that, in the case of K-Site, existed and then failed to travel into the earlier downstream papers examined — Note 001 records four accuracies and part of the thermal context returning only in 2025.

6. The report diagnoses its own instruments in public. The most striking passage in 286 pages concerns silicon diode SD76 during the baseline test (pp. 130, 140). The diodes’ “point-sense” mode — deliberately overpowering a diode and watching its thermal response to detect whether it sits in liquid or vapor — “did not work,” and the report says so. Worse, the apparent wet-to-dry transitions of the diodes can mislead: the report walks through a nine-hour trace showing an apparent transition at 15:38 that the RFMG data expose as having actually occurred nearly three hours earlier — “the apparent diode transitions can be in error by 3.7 percent or more” of full-scale fill, while the RFMG reading sat within 0.5 percent of full scale of the diode-location value. An instrument pathology, caught by a dissimilar instrument, published with timestamps.

That is what a referent built to travel looks like. Not error-free — instrumented, cross-checked, and candid.


III. The “almost”

Now the same reading, against the six criteria of Note 001. Because even here — and this is the point of auditing the best case — a modeler who sits down to validate against SHIIVER hits limits the report itself acknowledges.

The reference fill level is uncertain, and the uncertainty propagates. Because the point-sense mode failed, “it is not known where precisely the liquid interface was relative to the location of SD81” (p. 130) — and the report adds, with characteristic bluntness, that “it is moot whether the precise liquid interface location can be determined even if the point-sense mode worked.” The consequence is traced explicitly: the uncertain reference fill level feeds the constant C1 of the capacitance-probe correlation, and therefore every fill level computed from it. The report further notes that between atmospheric pressure and 172 kPa the liquid density changes 3.25 percent, warns that “Other errors will compound with this value,” and finds it “conceivable that the deviations shown in Figure 212 are reflective of the uncertainty in the fill level obtained from the analysis of the capacitance probe measurements.” Boiloff flow rates escape this particular chain only by anchoring on FM3 — via the stated assumption above.

I found no consolidated uncertainty budget. The report discusses uncertainty roughly two dozen times — voltage measurement where the DAQ gain “was defined too low,” the fill-level reference, wet/dry transitions, the heat flux sensors — but each discussion lives where the relevant instrument lives. What I did not find, searching the full text for uncertainty, accuracy, error, propagation, confidence, standard deviation and root-sum-square, and inspecting all 74 numbered tables, is a budget that propagates those individual instrument uncertainties into intervals for the headline results — heat load reduction 10 ± x percent; boiloff reduction 25 ± y percent. A modeler who wants u_D for a validation claim — the experimental uncertainty term that Note 001 showed is the first thing to stop travelling — must assemble it from fragments scattered across 286 pages. Substantial raw material is there. The assembled number is not — or, if it exists outside this report, the report does not say where — and whether the public record suffices to reconstruct every required contribution has not been demonstrated. (Re-verified at v1.0 over the corrected 74-table universe and the full report text: the report does quantify method-to-method dispersion — root-sum-square scatter between measured and calculated heat fluxes and loads, 7–69 % for fluxes and 10–60 % for loads depending on test and surface — but assembles no instrument-propagated uncertainty budget attached to a headline result. The rerun, with its search terms, is recorded in the verification log.)

Much of the time-history record is published only as plots. This was the finding that most surprised me. The report carries 74 numbered tables (46 in the main body, plus appendix series H.1–H.19, I.1–I.6 and J.1–J.3), but the bulk measurement record — skirt temperature profiles, vapor-cooling-line temperatures, per-test — lives in Appendices B through G as figures: “All plots (Figure B.1 to Figure B.63) are shown as a function of height along the forward skirt…” (p. 185). A modeler wanting those temperatures as numbers must digitize curves from a PDF, with the digitization uncertainty that adds. And I found no statement anywhere in the report that the underlying data files are archived in any public repository. I state that carefully: absence of a statement in the document I read, not evidence the archive does not exist. If it does exist, telling people where it is would cost a sentence.

The scorecard, then. Reading this case forced the criteria themselves to sharpen — an absolute “complete?” or “fully specified?” is not answerable in the abstract, only against an intended use:

Criterion (from Note 001, refined here)SHIIVER
1. Machine-readable measurement data?Partial — substantial numerical summaries across 74 numbered tables; much of the time-history record needed for model comparison exists only as plotted figures; no public machine-readable archive located, and the report does not say whether one exists
2. Boundary conditions?Strongly documented — chamber environment, MLI configuration, vapor-cooling flows; local heat-flux measurements exist but carry high absolute uncertainty (~50 %) and were compared against calculated system heat inputs — their stated uncertainty is not a budget for the system-level heat-input boundary condition. Whether the set is complete can only be judged against a specific model and quantity of interest
3. Measurement uncertainty?Strong but distributed — requirements specified in a standalone plan (Revision I, partly post-test — §II, practice 1); per-instrument candor throughout; no consolidated result-level budget located
4. Geometry?Substantially documented across report and test plan; exact computational reconstructability from the public record alone remains to be demonstrated
5. Fluid state?Analysis property basis explicit — parahydrogen, tabulated with references, the clearest single improvement over the 1992 generation. Whether the actual test fluid’s composition was independently measured is a separate question this note has not resolved
6. Does the quality information travel into downstream use?Open — scoped to downstream self-pressurization validation: the first such paper examined (2023) carried none of the context; see §IV. Other downstream use classes (propellant gauging among them) remain unexamined here

IV. Criterion six, measured for the first time

Scope: criterion 6 is measured here for one use class — downstream self-pressurization validation. A referent does not globally retain or lose its context; each quality item travels, or fails to, into each intended use.

Refinement recorded at v1.0 (Erratum 27): the 2023 paper’s thermal boundary condition is a separate Thermal Desktop model’s “preliminary estimates” of the heat loads — its own words, noting that lower dome values were later published in the experimental report — with a sensitivity run on dome heat. The materially missing uncertainty for that comparison is therefore the applied Thermal-Desktop boundary condition’s, which the paper does not quantify; the heat-flux sensors’ ~50 % figure is the documentary trace this criterion measures, not the paper’s applied input.

Earlier drafts of this note said the question was too early to tell: SHIIVER’s final report is from August 2021, and its downstream validation literature seemed young. That was an expiring negative of exactly the class Errata 16 and 17 record — because the first full-length downstream validation paper had been sitting on NTRS since January 2023: Kartuzova, Kassemi & Hauser, CFD Model Development of a Cryogenic Storage Tank Self-Pressurization in Normal Gravity and Validation against SHIIVER Experiment, AIAA SciTech 2023 (NTRS 20220017916, 21 pp., SHA-256 74ce93ac5493e9d9cbe42a6a7b3a88bae038087ed6dd227764799a92cd3a8838). It was obtained and read for v1.0, and criterion six stops being speculation here.

The audit is the one this note sets up: take the qualifications SHIIVER’s own documentation publishes, and ask which of them appear in the paper that validates against its data.

SHIIVER context item (this note, §II–III)Carried into the 2023 validation paper?
Boiloff analysis anchored on “(it is assumed that FM3 is accurate)”No — zero flowmeter mentions
Reference fill level uncertain; point-sense mode failedNo — zero mentions
Heat-flux sensors: uncertainties “on the order of 50%”No — zero mentions
Any experimental measurement uncertainty (± on a measured quantity)None stated
Source identifiedYes — the final report is cited

The paper’s headline is that predicted tank pressures are “within 4%” of the experimental ones — stated four times, with no experimental uncertainty against which 4 % could be judged. Its five uses of accuracy all refer to the model’s predictive accuracy, none to the data’s. This is, one referent later, the same shape Note 001 measured across the K-Site lineage: the data travelled; the qualifications did not.

One paper is one paper. The verdict for the referent stays open, the wider SHIIVER literature remains unsurveyed, and this note still makes no claim about the field. But too early to tell is no longer true, and the first measurement is in: the best-documented referent in this registry entered its first examined downstream use in the same documentary pattern this series found at K-Site: the data travelled, the relevant qualifications did not.

V. What SHIIVER should change about the audit

Reading the best case taught me two things that the audit — Technical Note 000, published alongside these notes — absorbs into its scoring scales.

First: “is uncertainty reported?” is too coarse a question. SHIIVER reports uncertainty in requirements, in fragments, and in candid asides — everything except a consolidated budget attached to the headline results. A referent can be simultaneously exemplary in uncertainty culture and unfinished in uncertainty bookkeeping. The audit distinguishes those on three levels: uncertainty required in advance; uncertainty characterized per instrument; uncertainty assembled per reported result.

Second: “are the data published?” needs a sharper edge. The honest categories are: machine-readable primary measurements; tabulated processed data; plotted-only data, with the digitization uncertainty that implies; an archive referenced but access-controlled; an archive publicly reachable; or not located. SHIIVER’s detailed time-history record remains substantially figure-based even though the report contains extensive tabulated summaries. The 1992 K-Site paper sits there too. In these two cases, thirty years apart, the instrumentation transformed, the candor deepened — and the format in which a modeler receives the actual numbers barely moved.

That last sentence — bounded to these two cases, until the audit says whether it generalizes — is, I think, the finding of this note.


VI. Corrections and additions wanted

If you worked on SHIIVER, on the eCryo project, or at the facilities involved, three things would be genuinely useful:

  • whether the underlying data files are archived anywhere a researcher can reach — the report twice states the data “have been archived” for future analysis, without saying where or how a researcher reaches them — the ask is the pointer, not the archive;
  • whether a consolidated uncertainty analysis exists beyond what the final report contains — an internal memo, a calibration package, an analysis that page limits cut;
  • cases where I have this wrong.

This note holds SHIIVER up as the counter-example — the referent whose caveats are recoverable because its authors wrote them down. If I have misread the documents that earn it that role, I want to know first.

Corrections will be credited. Where one changes a conclusion, it will be recorded as having changed it.

On verification. Every quotation in this note was checked verbatim against the source documents, and the report’s title page, authorship and identifiers were verified against its page images. Anything found wrong after publication produces a new version carrying a visible erratum — the superseded version remains preserved and citable — including when the correction weakens a conclusion already stated.



What follows this series

These notes are the documentary groundwork — published first so it can be checked first — for a committed quantitative study: reconstruct a defensible result-level experimental uncertainty (u_D) for one referent in this registry, combine it with the numerical and input uncertainties of a published comparison, covariances included, and report whether that comparison’s validation conclusion moves. In either direction: if nothing moves, that result publishes too. Until that study exists, everything here remains what Note 000 §IX declares — documentary findings whose engineering consequence is argued, not demonstrated. The study will appear as a new version under this series’ concept DOI, with these notes as its prior work.

Sources

A. Documents obtained and read for this note

  1. Johnson, W. L., Balasubramaniam, R., Hibbs, R., Zimmerli, G. A., Asipauskas, M., Bittinger, S., Dardano, C., Koci, F. D. — Demonstration of Multilayer Insulation, Vapor Cooling of Structure, and Mass Gauging for Large-Scale Upper Stages: Structural Heat Intercept, Insulation, and Vibration Evaluation Rig (SHIIVER) Final Report. NASA/TP-20205008233, August 2021, 286 pp. — https://ntrs.nasa.gov/citations/20205008233

  2. SHIIVER Test Plan, eCryo-PLN-0079 Rev. I, release date 9 January 2020. NTRS 20205003433, 53 pp. — https://ntrs.nasa.gov/api/citations/20205003433/downloads/eCryo-PLN-0079_SHIIVER_Test_Plan_RevI_2020-05-30_DLR%20(002).pdf

  3. Results of Use of Heat Flux Sensors on Liquid Hydrogen Tanks. NTRS 20210019122. — https://ntrs.nasa.gov/api/citations/20210019122/downloads/CEC_SHIIVER_htflux_paper_rev3.pdf

  4. Kartuzova, O., Kassemi, M., Hauser, D. — CFD Model Development of a Cryogenic Storage Tank Self-Pressurization in Normal Gravity and Validation against SHIIVER Experiment. Full conference paper, AIAA SciTech Forum, January 2023. NTRS 20220017916, 21 pp. SHA-256 74ce93ac5493e9d9cbe42a6a7b3a88bae038087ed6dd227764799a92cd3a8838. Read for §IV. — https://ntrs.nasa.gov/citations/20220017916

B. Framework

  1. ElarionX CPMS, Technical Note 000 — The Cryogenic Referent Registry (doi:10.5281/zenodo.21895568). Canonical wording of the six criteria, verdict scale, provenance and errata.
  2. ElarionX CPMS, Technical Note 001 — The Accuracies Travelled Back. The Warning Did Not. K-Site, 1992–2025 (doi:10.5281/zenodo.21895605). First publication of the criteria and of the u_D framing, which itself follows ASME V&V 20.

ElarionX CPMS develops and evaluates engineering models for cryogenic propulsion systems, and publishes what it finds — including about its own work. Competing interest: declared in Note 000 §IX. This note is exploratory: it states a reading and a method, not a validated result. Corrections and correspondence: Luis.emc2@elarionx.com

Version of record. The citable version of this note is the Zenodo deposit, doi:10.5281/zenodo.21895647. The text on this page is the same version; where they ever differ, the deposit governs. To cite the note across all future versions rather than this one, use the concept identifier doi:10.5281/zenodo.21895646.

Found an error? Corrections are wanted and will be credited. Where a correction changes a conclusion, the change is recorded as a change rather than edited away.

← All articles