Process Monitoring#
Class for ControlChart: robust control charts with a balance between CUSUM and Shewhart properties.
- process_improve.monitoring.control_charts.BIWEIGHT_RHO_CONSISTENCY = 3.266736#
Consistency constant for the bounded biweight rho with cutoff k = 2.52, c_k = 1 / E[rho_norm(Z)] for Z ~ N(0, 1), where rho_norm is the biweight rho normalised to a maximum of 1. Computed via numerical integration (scipy.integrate.quad of rho_norm(z) * phi(z); E = 0.3061160, so c_k = 3.266736). This makes E[rho(Z)] = 1, the condition for the scale estimates built from rho to be consistent for sigma on Gaussian data.
- process_improve.monitoring.control_charts.rho(x, k=2.52)[source]#
Bi-weight rho function.
Fixed cutoff of k=2.52 is from p 289 of the paper https://onlinelibrary.wiley.com/doi/abs/10.1002/for.1125
The multiplier is the consistency constant c_k (chosen so that
E[rho(Z)] = 1for standard-normal Z), NOT the cutoff k. The paper treats the two as separate constants; an earlier version of this code conflated them and used k = 2.52 as the multiplier, which made every scale estimate derived from rho a factorsqrt(2.52 * 0.30612) = 0.878too small, i.e. +/-3S control limits that were really +/-2.63 sigma (a ~3x inflation of the false-alarm rate).
- process_improve.monitoring.control_charts.psi(x, k=2.0)[source]#
Pre-clean based on the Huber psi function.
Can be interpreted as replacing unexpected high or low values by a more likely value. From p 288 of the paper https://onlinelibrary.wiley.com/doi/abs/10.1002/for.1125
- class process_improve.monitoring.control_charts.ControlChart(style='robust', variant='HW')[source]#
Bases:
objectCreate control chart instance objects.
- __init__(style='robust', variant='HW')[source]#
Create/initialize a control chart.
- Args: style (str, optional): Which style control chart to calculate. Defaults to “robust”.
Other choice is ‘regular’ (i.e. not-robust) calculations. User should then ensure that no outliers are present in the data.
- variant (str, optional): Only two variants are currently accepted:
'hw'(the default) and'xbar.no.subgroup'. Any other value, including'cusum', raisesValueErrorat construction time. The variant string is compared case-insensitively (it is normalised via.strip().lower()on assignment), so'HW','hw', and'Hw'are all equivalent.The default is a Holt-Winters (‘hw’) chart, with automatic determination of control chart parameters. This chart is a blend of infinite history (CUSUM) charts, and an instantaneous (no history taken into account) Shewhart chart. The exact blend is specified by parameters ld_1 (lambda 1) and ld_2 (lambda 2).
The other accepted variant is:
‘xbar.no.subgroup’ [Shewhart chart, with no subgroups]. In other words, each observation is independently plotted on the control chart.
A pure ‘cusum’ (CUmulative SUM) chart is a planned future variant but is not currently implemented; passing
variant='cusum'raisesValueError. The Holt-Winters (‘hw’) default already blends CUSUM-style infinite history with Shewhart-style instantaneous behaviour via its lambda parameters.
- calculate_limits(y, target=None, s=None, **kwargs)[source]#
Find for a given vector y, the control chart target and limits.
Works for both the Holt-Winters (‘hw’) and ‘xbar.no.subgroup’ variants.
For the Holt-Winters variant, when there are fewer than
min(20, max(10, np.ceil(0.10 * N)))
measurements (where N is the length of the input vector), the target and standard deviation are estimated directly from the data and any provided target / s are ignored for that small-sample case. Otherwise, if target and s are numeric, those values are used; if not, they are estimated.
- process_improve.monitoring.metrics.calculate_cpk(df, which_column, specifications=(nan, nan), trim_percentile=2.5)[source]#
Calculate the process capability, Cpk, near either the lower or the upper limit [will be automatically determined which].
Process capability, nearer the lower limit = (avg - lower_spec)/(3 x std deviation) Process capability, nearer the upper limit = (upper_spec - avg)/(3 x std deviation)
- Parameters:
df (pd.DataFrame) – Raw data, at least one column is numeric.
which_column (str) – Indicates which is the column of data that should be used for the Cpk calculation.
specifications (tuple of (lower, upper), optional) –
A 2-tuple
(lower_spec, upper_spec)of the lower and upper specification limits. Each element may be:a numeric value, when the specification is constant over time;
a string, interpreted as a column name in
dfwhose values give the per-row specification (use this when the specification changes over time);None, in which case the corresponding spec is estimated from the data usingtrim_percentile(a percentile-based robust limit).
Default is
(np.nan, np.nan), which treats both specs as numeric NaN and yields NaN for the corresponding side of the Cpk calculation.trim_percentile (float, optional) – Controls two things. (1) When a specification limit is missing,
trim_percentileis used as a percentile on the data (in percent) to estimate that limit: the lower spec is set tonp.nanpercentile(data, trim_percentile)and the upper spec tonp.nanpercentile(data, 100 - trim_percentile). Default2.5therefore yields the 2.5th and 97.5th percentiles. (2) Whentrim_percentile > 0the centre/spread used in the Cpk formula switch from mean/std to robust alternatives (median andSn); when 0 the classical mean/std are used.
- Returns:
A bunch with the following fields:
cpk: the Cpk value (the limiting, i.e. smaller, of the two sides).center: the center (mean or median) of the limiting side, i.e. of the distance-to-specification metric.spread: the spread (standard deviation or Sn) of the limiting side.rsd: the relative standard deviation OF THE DATA COLUMN, as a percentage:spread / center-of-the-data * 100. (Earlier versions divided by the distance-to-spec centre, which is not an RSD of anything and changed value when the spec limit moved.)
- Return type:
Notes
The spread is estimated from ALL the values in the column (the overall long-term variation), not from a within-subgroup range or moving range. Strictly this makes the statistic a performance index (Ppk-style) rather than a short-term capability index (Cpk with sigma-hat = Rbar/d2); on a stable process the two coincide, and on an unstable one this value is the more honest of the two.