Design of Experiments for Food Processes

RISK-BASED QUALITY AND INSPECTION

Experiment to understand, optimise and confirm

Turn process questions into evidence: select factors, recognise interactions, account for variability and confirm conditions that can improve food quality, safety and efficiency.

DECISIONS SUPPORTED BY EVIDENCE

What does design of experiments offer?

Design of experiments (DOE) deliberately changes input variables to estimate their effects on one or more measured responses. Unlike passive analysis of historical records, it generates information under planned conditions and helps distinguish factor effects, interactions and random variation.

Understand

Identify which factors drive meaningful changes in yield, texture, colour, shelf life or safety-related responses.

Optimise

Find combinations that bring several responses closer to their objectives while respecting technical and commercial constraints.

Confirm

Use new, independent runs to check whether the predicted improvement is reproducible rather than an isolated result.

BEYOND ONE FACTOR AT A TIME

Why interactions matter

Changing one factor at a time (OFAT) may seem straightforward, but a conventional OFAT sequence cannot estimate interactions. In food processes, the influence of temperature may depend on time, pH, moisture or formulation.

One factor at a time

  • Explores only a limited part of the experimental region.
  • Can confound changes with time-related drift.
  • Does not estimate interactions in a conventional OFAT sequence.
  • May lead to a suboptimal combination.

Factorial design

  • Varies factors in a coordinated pattern.
  • Estimates main effects and estimable interactions.
  • Uses each run to learn about several factors.
  • Supports a model that can be tested and confirmed.

An interaction occurs when the effect of one factor depends on the level of another. It may reflect a genuine process mechanism. A main effect averaged across other factors can therefore hide important differences.

ESSENTIAL VOCABULARY

The building blocks of a defensible experiment

Factor and level

A factor is an input deliberately varied; its levels are the conditions tested. Factors may be quantitative, such as temperature, or categorical, such as packaging type.

Response

The measured outcome: yield, moisture, firmness, microbial count, nutrient loss, sensory acceptance or another defined indicator.

Experimental unit

The smallest unit independently assigned a treatment. Distinguishing it from a subsample prevents pseudoreplication.

Replication

Independent repetition of a treatment combination, used to estimate experimental error and assess reproducibility. Repeated readings of one sample are not independent experimental replicates.

Randomisation

Random assignment or run order helps prevent systematic association between treatments and uncontrolled influences such as time, position or operator.

Blocking

Group similar experimental units by a known nuisance source, such as day, raw-material batch or equipment. Include the block structure in the analysis.

A PRACTICAL WORKFLOW

From the question to confirmation

01 · Define the objective

Specify the decision, responses and smallest technically meaningful change.

02 · Choose factors and levels

Set safe, feasible ranges, constraints and conditions to keep constant.

03 · Select the design

Match the run count, replication, blocks and statistical power to the objective.

04 · Execute and record

Follow the planned allocation and run order, measurement procedure and traceability requirements.

05 · Build the model

Plot the data, estimate effects, examine residuals and quantify uncertainty.

06 · Interpret

Distinguish statistical significance from technological importance and consider interactions.

07 · Confirm

Run fresh experiments at selected settings and compare observed responses with predictions.

08 · Transfer and monitor

Translate supported findings into operating conditions and monitor the process with SPC.

SELECTING THE DESIGN

Different designs for different questions

A starting point—not a substitute for planning
ObjectiveTypical designWhat it offersKey caution
Compare treatmentsCompletely randomisedA simple structure for estimating treatment differences.Best suited to reasonably homogeneous units; account for important heterogeneity.
Account for known nuisance variationRandomised blocksSeparate variation associated with batch, day, line or operator.Do not confuse the block with the factor of interest.
Study a small set of factorsFull factorialEstimate main effects and interactions.The required run count grows rapidly.
Screen many factorsFractional factorialIdentify influential factors with fewer runs.Check aliasing: some effects cannot be separated without assumptions or additional runs.
Check for curvatureTwo-level factorial with centre pointsTest whether a first-order description is inadequate for quantitative factors.Centre points alone do not identify all quadratic terms.
Optimise a processResponse surfaceModel curvature and find a promising operating region.Usually focus first on a smaller set of important factors.
Optimise a formulationMixture designModel component proportions constrained to a total.Do not treat all proportions as ordinary independent factors.

Some factors, such as oven temperature, may be hard to change between runs. If randomisation is restricted, use a design and analysis that reflect this—for example, a suitable split-plot structure—rather than pretending that every run was independently randomised.

ANALYSIS AND INTERPRETATION

ANOVA is evidence, not the whole decision

Analysis of variance (ANOVA) partitions variation into model terms and residual variation and supports tests of effects. A p-value alone does not establish technological importance, economic value or reproducibility.

Check before concluding

  • Plots of responses, main effects and interactions.
  • Residual patterns, variance assumptions and influential observations.
  • Effect magnitudes and confidence intervals.
  • Model hierarchy when interactions are retained.
  • Measurement-system adequacy and experimental-unit independence.

Avoid weak conclusions

  • “Not significant” does not demonstrate no effect.
  • More replicates cannot fix a biased design.
  • Averaging can conceal an interaction.
  • A good fit does not justify extrapolation.
  • Optimisation without confirmation leaves the decision incomplete.
A two-factor model

Y = β₀ + β₁A + β₂B + β₁₂AB + ε

A and B are coded factor levels, β terms are model coefficients and ε represents random error. AB is the interaction term. With −1/+1 coding in a balanced two-level factorial, the conventional factor effect is twice its fitted coefficient; the coefficient is not itself the high-minus-low effect.

A full 2k factorial contains 2k combinations before replication or centre points. One observation per combination with all factorial terms fitted leaves no residual degrees of freedom. Do not report routine significance tests from such a saturated model without a defensible error estimate and explicit assumptions.

ILLUSTRATIVE FOOD-PROCESS CASE

Initial optimisation of biscuit baking

A team wants to reduce final moisture without excessive colour development or loss of yield. It selects three factors at two levels and adds centre points to investigate possible curvature.

Illustrative factor ranges from the Spanish guide
FactorLow level (−)High level (+)Rationale
A · Temperature170 °C190 °CAn illustrative oven operating range.
B · Time8 min12 minContrasting conditions for moisture removal.
C · Fat content16%20%Potential effects on texture, spread and heat transfer.

The full 2³ factorial has eight treatment combinations. Replicates and centre-point runs are additional. The centre corresponds to 180 °C, 10 min and 18% fat. Plan independent replication and randomise the run order within any justified blocks.

Before execution, define the basis of the fat percentage and how the remaining formulation changes. If the question concerns component proportions jointly constrained to a total, consider a mixture–process design. Define the experimental unit at the treatment-assignment level; several biscuits from one batch or oven run are not automatically independent replicates.

Responses include moisture content, colour difference ΔE, yield and fracture force. Define the methods, sampling positions and acceptable response ranges before the trial.

Hypothetical findings

  • Temperature and time reduce moisture.
  • The temperature × time interaction strongly affects colour.
  • Fat alters texture and partly moderates yield loss.
  • Centre points suggest curvature in the colour response.

A reasonable next decision

Do not simply select the hottest, longest treatment. Explore an intermediate region using a suitable response-surface design, set acceptable limits for all responses, and conduct independent confirmation runs.

These values and findings are educational illustrations, not measured trial results or recommended settings. A real experimental region must account for the food, process safety, equipment capability and applicable requirements. Moisture reduction or acceptable colour alone does not validate a safety outcome.

AN INTEGRATED QUALITY SYSTEM

When to use DOE, SPC, sampling or full inspection

Choose the method according to the decision
QuestionMain methodTypical use
Which inputs influence the response, and which conditions are promising?Design of experimentsDevelopment and improvement.
Does the process remain stable over time?Statistical process controlRoutine operation.
Should a lot be accepted under a defined decision-risk framework?Acceptance samplingReceipt or release decisions.
Can each unit be evaluated quickly, non-destructively and reliably?100% inspectionWhere the technology and risk assessment justify it.

The methods complement each other. DOE develops process knowledge; SPC monitors stability; acceptance sampling supports lot decisions; and full inspection can detect nonconforming units for the characteristics tested. None automatically replaces the others or guarantees the absence of hazards.

COMMON PITFALLS

What to avoid

During planning and execution

  • Starting without a clear operational question.
  • Choosing unsafe or irrelevant ranges.
  • Confusing subsamples with independent replicates.
  • Ignoring randomisation or its restrictions.
  • Overlooking batches, days or operators as nuisance sources.

During interpretation and transfer

  • Removing interactions just to simplify the model.
  • Improving one response at the expense of another.
  • Extrapolating beyond the experimental region.
  • Selecting settings only from p-values.
  • Omitting independent confirmation runs.

SOURCES AND FURTHER READING

Technical references

  1. NIST/SEMATECH — Process Improvement: experimental design, analysis and confirmation.
  2. NIST — Choosing an experimental design: randomisation, blocking, factorial and fractional factorial designs.
  3. NIST — How do you select an experimental design?: matching objectives to design families.
  4. NIST — Two-level full factorial designs: combinations, coding and replication.
  5. NIST — How to model DOE data: model structure, aliasing and validation.
  6. FDA/ICH — Q8, Q9 and Q10: Points to Consider: supplementary pharmaceutical-sector context on experimentation, models and lifecycle control; not a food-industry regulatory requirement.

An educational QualiFood adaptation, not an official translation or substitute for process-specific validation and professional judgement.

From experimental design to practice

Build a 2² or 2³ factorial design, define factors and levels, enter responses and explore main effects and interactions. The interactive explorer is currently available in Spanish.