A fleet profile is a hypothesis. A cohort result is evidence.
Picture the change window. The fleet console shows a long block of S21-class machines. The labels look close enough. One profile is ready. Select all would be fast.
But the rack does not operate as a spreadsheet row.
Some units may sit on another control-board route. Some may sit on a different PSU or power-feed group. Air at the end of a row may not match air at the center. A replacement hashboard, a history of resets, or a changed firmware build can put one machine on a different response path from its neighbor.
None of that proves a problem is common. It proves that the label alone cannot answer the rollout question.
The practical move is not to abandon fleet control. It is to define the fleet you actually have: coherent cohorts, one controlled canary per cohort, one untouched peer or control group, a declared observation window, and a decision that can split the rollout when the distribution says it should split.
- Do not use one model label as the only rollout key.
- Build cohorts from exact hardware route, power-feed context, cooling zone, build state, machine condition, and observed response.
- Keep configured firmware controls separate from measured wall power and the site-approved wall-power ceiling.
- Compare one canary per coherent cohort with an untouched peer or control group.
- Report the median and the distribution tail or bands, not only the mean.
- Decide GO, HOLD, SPLIT, or STOP for each cohort, not automatically for the whole fleet.
Why the fleet average is not enough
The mean is useful. Add the readings and divide by the number of machines. It gives one center for the group.
The problem begins when that center is treated as the group.
Imagine a cohort where most machines hold their operating point while a smaller tail repeatedly resets. A mean hashrate can remain visually calm because the stable majority dilutes the weak tail. The same can happen when a few machines carry higher wall power, wider thermal oscillation, more errors, or less accepted work than the center.
The mean did not lie. It answered a narrower question than the operator asked.
NIST describes a distribution through location, spread, and shape. Its measures-of-location guidance notes that extreme tail values can distort a mean, while the median is based on rank and is less affected by those extremes. Percentiles place ordered observations into relative positions. These are general statistical tools, not mining thresholds.
For a fleet decision, report at least three views:
- Center: median, with the mean retained when it is useful.
- Spread: a band, interquartile range, or another declared view appropriate to cohort size.
- Tail: the machines closest to the cohort's risk direction, such as low accepted work, high wall power, repeated resets, or unstable thermals.
Do not assign a universal percentile. A large cohort may support percentile bands. A small cohort may be clearer as a sorted list showing every machine. The operator must declare the method, direction, and window before using the result.
The average describes the fleet you can calculate. The tail shows the machines you still have to operate.
Follow the signal chain before forming the cohort
Each step is a question. The sequence does not prove causation.
Nameplate
The model label starts inventory. It does not prove the control board, firmware artifact, PSU context, condition, or response. The live VNISH control-board map shows why exact routes matter across the supported catalog. It does not imply that every base S21 has multiple board routes.
For the exact base S21, BITMAIN lists 200 TH/s and 3,500 W power on wall at 25 degrees C. Those values belong to that exact model and stated condition. They are not a field threshold for an individual unit, an adjacent S21 model, or a fleet cohort.
Exact hardware route
Record the exact model, control board, current build, target build, and supported install route. Keep artifact verification in the existing S21 identity guide rather than repeating it here.
The current VNISH 1.3.5 release notes say presets were reworked for several named models including S21 and keep installation, autotune, and final stabilization separate. That supports recording the build and stage. It does not predict a cohort result.
Power and cooling zone
Tag the PDU branch or other site-approved power-feed group, PSU context, and the cooling or ambient zone that could make machines comparable or non-comparable. The zone must reflect the site's actual layout and measurements, not a generic label copied across buildings.
Configured control
Record the exact profile, limit, or setting label and the build that exposed it. A configured firmware limit is a control setting. It is not measured wall power and it is not the site-approved wall-power ceiling.
Observed wall and machine telemetry
Name the wall-power instrument and measurement boundary. Preserve the firmware-reported watts label separately. The VNISH platform page describes requested and observed operating points, watts, measured efficiency, temperatures, fan response, board state, errors, and active profile as operational signals. Its displayed demo values are examples, not fleet evidence or a performance promise.
Pool-accepted work
Pool reporting is a separate observation from a target or device-side hashrate. The Stratum V2 specification distinguishes submitted work, successful acceptance responses, accepted-submit counts, and the sum of acknowledged share difficulty. Use the pool's defined fields consistently and keep the endpoint and window aligned across the comparison.
Stability and decision
Combine uptime, resets, errors, missing boards or chips, thermal behavior, accepted work, and the site's electrical and cooling rules. No single telemetry tile clears a cohort.
The Cohort Card
Use one row per machine. The card is deliberately wider than a firmware dashboard because the rollout decision crosses hardware, site, firmware, wall measurement, pool outcome, and ownership.
| Field | Record | Why it belongs |
|---|---|---|
| Model | Exact model string | Starts identity without pretending to finish it |
| Control-board route | Exact board family and supported route | Keeps incompatible routes out of one wave |
| PSU and power-feed group | PSU context, PDU branch or approved site grouping | Preserves shared electrical context |
| Cooling or ambient zone | Named inlet, row, duct, loop, or site-defined zone | Keeps unlike thermal exposure visible |
| Firmware and build route | Current build, target build, install route | Makes the change reproducible |
| Configured profile or limit | Exact label, value, build, time, operator | Records the firmware control without calling it wall power |
| Site wall-power ceiling | Approved ceiling, measurement boundary, owner | Separates a physical site constraint from a setting |
| Measured wall power | Instrument, boundary, value or range, timestamps | Records demand at the named physical boundary |
| Reported watts | Exact interface label, value or range, timestamps | Preserves the device signal without relabeling it |
| Observed hashrate | Exact source, label, and window | Records machine-side observation |
| Accepted pool work | Pool, endpoint, field, value, and window | Records recognized work separately |
| Rejects or stales | Pool fields when available | Adds delivery quality without inventing missing data |
| Thermal behavior | Exact labels, range, trend, oscillation, fan response | Shows behavior across the window, not one peak |
| Errors and resets | Counts, types, timestamps, missing hardware | Makes instability visible |
| Observation window | Start, end, timezone, operational cycles, interruptions | Makes comparisons reproducible |
| Decision owner | Named role and GO, HOLD, SPLIT, STOP authority | Prevents an ownerless rollout |
If a field is unavailable, mark it unavailable. Do not silently replace it with another signal.
Five cohort lenses, not five universal categories
The useful cohort depends on the change being tested. The following are operator-defined lenses, not a universal classification system.
Route cohort
Group machines that share exact model, control board, current build, target build, and install route. This is the minimum compatibility wave described by the official VNISH fleet page.
Power-delivery cohort
Group machines only when the PSU context, feed, measurement boundary, and site-approved ceiling make them comparable for the question at hand. A firmware setting does not replace this grouping.
Cooling-zone cohort
Use the site's measured airflow, inlet, containment, hydro loop, or ambient layout. Two adjacent asset numbers may sit in different thermal conditions. Two distant units may share a controlled condition. The site defines the zone from evidence.
Use the S21 thermal baseline guide when the cohort needs deeper inlet, airflow, and thermal context.
Condition cohort
Separate clean candidates from units with unresolved resets, board errors, fan issues, repairs, sensor anomalies, or maintenance flags. Do not use an unhealthy machine as the sole representative of a healthy cohort, or hide it inside the average.
If issues persist after controlled observation, use the tune, repair, or replace guide to structure weak-tail triage.
Response cohort
Create this only after the canary evidence exists. Machines that start in one pre-change cohort may split into stable-center, review-tail, or another operator-defined response band. A response cohort is an observation, not a permanent identity.
Do not create one giant label containing every possible tag. Start with the variables that can plausibly change the decision. Preserve all fields in the card, then split only when the evidence or site design justifies it.
A practical cohort rollout protocol
- Define one change and one decision ownerName the exact build or profile question. Record who can approve the site wall-power ceiling, who owns the firmware change, and who can call STOP.
- Inventory and tag the fleetUse direct records for model, board route, build, power-feed group, cooling zone, and known condition. The official VNISH data catalog can resolve build-route records. It does not contain your site's electrical, cooling, or condition data.
- Form coherent pre-change cohortsChoose the smallest set of tags that matters to this change. Publish the cohort definition internally before selecting results. Machines with missing critical evidence stay out of the wave.
- Keep an untouched peer or control groupFor each coherent cohort, leave a comparable machine or control group on its existing build and profile. Choose one canary from the cohort. Do not use a unit with unresolved faults merely because it is convenient.
- Declare the observation windowChoose a window that covers the operational cycles relevant to the site, which may include thermal, load, network, pool, or staffing changes. Record start, end, timezone, warm-up treatment, interruptions, and site events. This article sets no universal duration. For deeper site-event and curtailment context, use the flexible-load operating guide.
- Apply, observe, and preserve the chainApply only the exact supported change. Keep the configured control, site-approved wall ceiling, measured wall power, machine telemetry, accepted pool work, errors, resets, thermals, and stability as separate fields. Preserve time-series evidence when available.
- Compare the center, spread, and tailUse the same method and aligned windows for canary and control. Report the median, a declared spread or band, and the risk-direction tail. For small cohorts, show every sorted unit. Do not pool unlike cohorts to make the result look smoother.
- Decide, then stage the next waveUse the card below. Record the evidence and owner. A cohort may move while another holds or splits. The whole point is to avoid forcing one answer onto machines that did not produce one answer.
GO, HOLD, SPLIT, or STOP
The coherent cohort met its prewritten, site-specific criteria through the declared window.
Advance only that cohort to the next controlled wave.Evidence is incomplete, the window is unrepresentative, or the comparison is not yet valid.
Keep the cohort unchanged and resolve the evidence gap.The distribution or site context shows two or more coherent response groups.
Create new cohort definitions and canaries before expansion.A prewritten electrical, thermal, hardware, accepted-work, error, reset, or stability condition was reached.
Contain the change, preserve evidence, and follow qualified support or recovery procedure.GO does not mean the entire fleet is cleared. HOLD does not mean the profile failed. SPLIT does not let weak-tail machines disappear from the report. STOP does not create one universal threshold for another site.
SPLIT is the decision most fleet averages cannot express. It lets the stable center remain evidence without sacrificing the tail to the center.
What this does not prove
A cohort comparison does not prove that firmware caused every observed change. Site events, machine condition, instrumentation, network path, pool behavior, maintenance, and unrecorded differences can also move the result.
One observation window does not forecast another season, difficulty period, site configuration, or hardware age. A cohort definition that works for one change may be wrong for the next.
No percentile, band, or median is a universal pass line. No result transfers automatically to another model, control board, PSU, cooling system, electrical design, pool, firmware build, or facility.
Firmware and profile changes can affect power draw, temperature, stability, component stress, hardware life, and manufacturer warranty or support conditions. No hashrate, efficiency, wall-power, temperature, uptime, savings, recovery, revenue, or financial result is guaranteed. Qualified personnel must approve electrical limits and work.
The next useful action
Save the Cohort Card. Choose one pending fleet change. Tag the machines by exact route, power-feed context, cooling zone, build, and known condition. Select one canary and one untouched peer per coherent cohort. Declare the window and the decision owner before the profile changes.
Then follow the evidence path: verify device identity, build the wall-power envelope, and keep machine and pool measurements distinct.
VNISH Global maintains the official build routes, release records, platform guidance, and fleet-governance tools. Use them to define the change. Use your own measured cohorts to decide whether it scales.
Compare this guide with six dated, source-linked VNISH field reports before turning one operator result into a fleet assumption. Open the independent field reports.
Sources and working references
- VNISH Global, Operational control in one interface. Signal taxonomy and operating context.
- VNISH Global, Governed deployment for ASIC fleets. Pilot cohorts, compatibility waves, decision ownership and rollout discipline.
- VNISH Global, Firmware catalog as a dataset. Exact model, control-board route, build, release, file and install records.
- VNISH Global, Control boards. AML, XIL, CV and BB route families.
- VNISH Global, Release 1.3.5. Current release context and separate install, autotune and stabilization stages.
- BITMAIN, S21 Specification. Exact base S21 reference values and stated conditions.
- Stratum V2 Mining Protocol Specification. Submitted work and accepted responses.
- NIST/SEMATECH e-Handbook, Distribution. Location, spread and shape.
- NIST/SEMATECH e-Handbook, Measures of Location. Mean, median and extreme values.
- NIST/SEMATECH e-Handbook, Percentiles. Ordered observations and relative positions.
VNISH Global is part of the VNISH firmware ecosystem. This guide is editorial material, not electrical, safety, warranty, investment, tax, legal, or financial advice. Official hardware documentation, live build routes, site procedures, and qualified personnel take precedence.