Sensor reading error: cause tree and tests

A spike on the chart can come from the structure, the sensor, the cable, time, the formula, or the reference. This procedure leads from the cheapest data tests to field verification.

Direct answer

When a sensor shows a suspicious value, start by preserving evidence and checking the data layer: time, completeness, the raw frame, and the calculations. Then inspect the logger, cable, power supply, and environmental influences. Only after those tests decide whether the structure has changed. Do not zero the channel or correct the formula before securing the state.

In brief

  • A suspicious reading is a signal that requires inspection, not a final diagnosis of a sensor or asset failure.
  • First preserve the raw measurement, time, server response, configuration, reference, and neighboring data.
  • Diagnose four layers in order: data, measurement chain, configuration, and environment, and only then the structure.
  • The shape of the error narrows the search, but it does not prove the cause on its own.
  • If a real change cannot be ruled out quickly, use the safe action from the alarm procedure.

A suspicious reading is the start of an investigation

A sensor reading error is a value that does not match the true value of the measured quantity at that moment, not simply a number that differs from what was expected. This definition prevents a common shortcut in thinking: "the chart looks odd, so the sensor has failed." The structure may behave differently from the model. Ground conditions may change faster than the schedule assumed. Unusual does not mean wrong.

ISO 18674-1 treats a monitoring system as a whole that includes the instrument, transmission, acquisition, and auxiliary components. It also notes that the parameters of the whole system do not have to be identical to the parameters of its parts. That matters: a healthy transducer does not guarantee a correct result after the cable, logger, time, formula, and reference have all been applied.

The first decision after an anomaly is evidential. Before anyone changes a coefficient, zero reading, or wiring, the state must be preserved. Otherwise the repair may erase the clue. An hour later the panel will look "clean", but it will no longer be possible to show whether the cause was configuration, moisture in the connector, or real movement.

The second decision concerns safety. Data diagnostics must not delay the action defined for an alarm. If the value exceeds the set threshold and quick tests do not show an obvious error, it should be treated as potentially real. The measurement procedure and the response procedure run in parallel.

For you, that means one thing: in the first minutes, the goal is not to prove sensor failure. The goal is to preserve the ability to decide the cause without exposing the asset to risk.

The four-layer cause tree

It is worth moving through the tree from fast, reversible tests to field work that requires site access. Each "I do not know" answer moves to the next layer. You must not jump straight to changing the threshold, because that is not a cause test.

Layer Typical question Evidence Next step
1. Data and time Has the sample arrived and does it have the correct timestamp? Raw frame, log, cadence Compare the record with the chart
2. Measurement chain Are the sensor, cable, and logger working? Visual inspection, resistance, control reading Test components from the end of the chain
3. Calculation and environment Did the formula, reference, or temperature change the result? Configuration before/after, contextual series Recalculate on a copy and compare
4. Structure Is there an independent trace of physical change? Adjacent points, independent measurement, inspection Engineering assessment and procedure

Layer 1: data, time, and transmission

Check whether the raw frame contains the same value shown on the panel. Compare the measurement time with the reception time. Look for duplicates, gaps, late samples, and changes in order. If several device series shifted at the same time, the clock or a shared transmission element may be at fault.

Missing data is a separate state. You cannot infer from it that the measured quantity was stable. A flat line, in turn, may mean a stable asset, a value frozen in the logger, or repetition of the last sample. To decide, you need the raw input and freshness information.

Layer 2: sensor, cable, power, and logger

The manufacturer's documentation for the specific instrument takes priority here. For example, the GEOKON manual for vibrating wire piezometers, in the case of unstable readings, instructs the user to check shielding, interference from motors and generators, overload damage, and short circuits. When there is no reading, it points to checking the logger with another sensor and measuring cable resistance. These are tests for a specific device family, not a recipe for every sensor. The detailed test sequence for this technology is described in vibrating wire sensor diagnostics.

The safest approach is to test from common elements. If all channels on one logger jump, start with power supply, grounding, configuration, and the link. If the problem affects only one channel, narrow the investigation to the input, cable, and sensor. Do not open a sealed transducer in the field if the manufacturer forbids it.

After excluding the measurement chain: calculations, environment, and structure

If the data, time, and basic measurement chain tests do not explain the deviation, move to the calculation method and to independent signs of change in the asset. Continue moving from reversible evidence to actions that could change the series or configuration.

Layer 3: equation, reference, and external influences

The raw value may be correct while the engineering result is wrong. A changed coefficient, a mistaken unit, an incorrect sign, a wrong reference, or missing temperature compensation is enough. Compare the configuration before and after the change. Recalculate several samples manually or in a controlled environment, without overwriting production data.

Temperature deserves its own chart. Daily correlation may point to an environmental influence, but it does not prove that the entire change came from temperature. The assessment method is described in the article temperature compensation or structural change.

Layer 4: the asset

At the end, look for independent confirmation: movement of neighboring points, a change in another measurement type, a field observation, a construction event, or external conditions matching the time of the anomaly. USACE recommends combining data with knowledge of asset behavior and comparing it with other instruments. A single chart is not enough.

The shape of the series gives a lead, not a verdict

A chart helps set the order of tests. It is not a cause reader. The same shape can arise in several ways, so the table shows the first test, not a final diagnosis.

Symptom First check Possible physical layer
Single spike and return Time, interference, connection, field event Short impulse or load
Spike and new level Reference, formula, logger restart, damage Permanent displacement or state change
Growing noise Shielding, moisture, power supply, range Variable process or vibration
Slow drift Temperature, sensor stability, reference Slow structural or ground change
Flat line Sample repetition, resolution, input lock Real stability
Many channels at once Shared clock, logger, power supply, calculation Shared load or temperature

A particularly deceptive case is a spike after a configuration change. If the equation or reference was changed between two samples, the chart may show a discontinuity without any change in the raw reading. That is why a log of the values before and after is more important than the comment "settings were corrected".

The second difficult case is the apparent agreement of neighboring channels. Two sensors may share the same logger, power supply, temperature, or incorrect formula. Correlation increases the credibility of a physical change only when you understand the shared technical dependencies.

Before the service visit, prepare a short hypothesis sheet. For each hypothesis, write the test, the expected result if confirmed, and the result that would weaken it. This way the technician does not go out to "check the sensor", but to measure resistance, compare the reading on another input, or inspect a specific section of cable. Such a sheet limits random adjustments and lets the visit end with evidence, not impressions.

Illustrative example: pore pressure spike

Assume, for illustration, that a piezometer channel shows an 18 kPa increase at 08:15, and the alarm threshold is exceeded. This is not a real deployment or a threshold to copy. The purpose of the example is to show the sequence of evidence.

The operator confirms the alarm and preserves the raw frame. Device time matches the server, and the samples before and after are present. The raw value contains the spike, so a charting error is ruled out. The neighboring piezometer does not react, but it lies in a different soil layer. Temperature is stable. The construction log records the start of pumping 20 minutes earlier.

The service team checks the logger on another input, the connector condition, and the cable according to the manufacturer's instructions. No short circuit or instability is found. The control reading repeats the result. The engineer compares the change with the expected response of the hydraulic system and checks the pump status and water level at an independent point. Only this set of evidence allows an assessment of whether the result is a soil response, a local installation issue, or instrument damage.

The bad version of this story is shorter: someone labels the spike a "sensor error", sets a new reference, and brings back the green status. Then both the alarm and the ability to check the cause disappear. If the movement was real, the team has just turned off its warning signal.

The procedure should also contain a condition for stopping remote diagnostics. It may be a second independent exceedance, a rapid rise, visible damage, loss of multiple critical channels, or the inability to confirm the data within the required time. Then the priority becomes the action defined for the asset: securing the area, limiting works, or carrying out a field inspection. Continued searching for an error in the panel must not become an excuse to postpone that decision.

Evidence checklist before changing the configuration

Not every project has a full diagnostic package. Even so, a minimum can be defined before the monitoring system is commissioned. The list below is suitable for a service procedure.

  • [ ] Raw frame with value and unambiguous time.
  • [ ] Server response and technical reception time.
  • [ ] Data range before and after the event, without smoothing.
  • [ ] Raw value, calculated result, unit, and formula used.
  • [ ] Current reference and the date it was established.
  • [ ] Device and project configuration before the change.
  • [ ] Temperature data and dependent channels.
  • [ ] Communication, power supply, and other logger input status.
  • [ ] Photos of connectors, cable, and installation location, if an inspection was performed.
  • [ ] Control measurement result and the tool used.
  • [ ] Work, weather, or operating log from the same time.
  • [ ] The person who made the decision and the reason for the change.

If the raw communication log has limited retention, download it immediately. Otherwise, after a few days only the processed series will remain, and some questions will go unanswered. For broader quality analysis, seven metrics for measurement data SLA can help.

What it looks like in Inclify

Inclify lets you plot multiple series on one chart, for example the measurement result and temperature, including on two axes. For exchanges with devices, the platform keeps a communication log containing the compressed raw request and response plus metadata. An administrator can view and download their raw content. The default log retention is seven days and it is configurable, so after an incident it is not worth delaying evidence preservation.

Engineering values are generated from project equations. The platform validates syntax, dependencies, and cycles, and the reference can be selected from an existing measurement or set manually. Changes to the project, device, and other audited configurations leave before and after values. Audit, however, is not a function that automatically restores the previous state.

A separate NO_DATA alarm detects the lack of fresh measurements for a device or channel. It helps distinguish a quiet chart from interrupted transmission. The platform does not automatically decide whether an unusual value comes from the sensor or the structure. It provides the traces needed for the next tests.

The greatest value comes from placing these traces on one timeline before changing the reference, equation, or settings responsible for interpreting the series.

Procedure limitations

The cause tree organizes the work, but it does not replace the manufacturer's instructions or the experience of the person examining the asset. A resistance test suitable for one sensor type may be meaningless or unsafe for another. Always verify the documentation for the specific model and the warranty conditions.

Not every cause can be confirmed after the fact. A short disturbance may disappear before the service visit. A damaged cable may work depending on humidity. Structural movement may be local and invisible at neighboring points. That is why redundancy, field verification, and good records are part of the monitoring design, not an add-on for failure time.

The platform does not replace independent measurement. If the consequences of the decision are serious, verify the reading by another method and with the appropriate engineer involved. Change the reference or formula only after the evidence has been recorded, the cause has been approved, and the impact on history has been defined. When the discrepancy grows, also use the procedure described in the text on sensor calibration drift.

FAQ

Does a single spike mean the sensor is damaged?

No. It may result from interference, a loose connection, a timing error, a short load, or a real event. Preserve the raw sample and the data on both sides of the spike, verify shared channels and work context. If the value exceeds the threshold, apply the response procedure until the issue is clarified. Do not erase the point before securing evidence.

When may a new reference be set?

When it is a justified engineering decision and the old state, the cause of the change, and the impact on history have been recorded. A new reference is not a tool for clearing an alarm. It should have a date, conditions, an approving person, and a link to the work stage or instrument replacement. The control result should be attached to the event documentation.

How do you distinguish a sensor error from temperature influence?

Set the raw value, temperature, compensated result, and comparison channels on the same time axis. Check whether the relationship repeats over several cycles and whether it changed after intervention in the installation. Correlation alone does not prove causation. The check must be based on a physical model and independent data.

Is missing data the same as a faulty reading?

No. Missing data means that no new sample exists in the expected window. A faulty reading is a sample that does not represent the measured quantity correctly. Both cases require different tests and different messages, although they may have a shared cause, such as a damaged cable or power supply.

How long should raw incident logs be kept?

For as long as required by the project procedure, the contract, and the evidential purpose. The default retention of the technical communication log in Inclify is seven days, but it is configurable. After an event, the administrator should download the relevant frames, responses, and metadata immediately rather than rely on indefinite availability.

Can statistical analysis point to the cause?

It can detect an unusual point, a spike, drift, or correlation and help set the order of investigations. It does not prove the physical mechanism, however. The cause is confirmed by checking the data, measurement chain, configuration, environmental conditions, and the asset. Final interpretation belongs to the person responsible for the technical assessment.

Sources and further reading

  1. International Organization for Standardization, ISO 18674-1:2015, Geotechnical monitoring by field instrumentation, general rules.
  2. U.S. Army Corps of Engineers, EM 1110-2-1908: Instrumentation of Embankment Dams and Levees.
  3. National Institute of Standards and Technology, NIST SP 1800-10: Protecting Information and System Integrity in Industrial Control System Environments.
  4. GEOKON, Model 4500 Series Vibrating Wire Piezometer, Troubleshooting.
  5. U.S. Army Corps of Engineers, ER 1110-2-1156: Safety of Dams, Policy and Procedures.

What next

Choose one real incident from the last few months and identify which pieces of evidence could no longer be recovered today. That is the best test of the procedure. Arrange a call with the Inclify team if you want to run the cause tree on data from one asset and assess whether logs, references, and audit records give service enough material.

Keep reading

Related articles

All articles

Vibrating wire sensors: why they are still the standard

A vibrating wire sensor measures the frequency of a tensioned wire, not voltage or resistance. That is why it can still deliver reliable readings after decades in concrete. Learn the operating principle, temperature compensation, drift, calibration debt, the 20-year case, and comparison with MEMS and fibre optics.

Read more →

Monitoring a structure? Book a demo

We will show the platform using an asset similar to yours and discuss where the measurement programme should start. No obligation.