The Upgrade Found Problems That Had Nothing to Do With the Upgrade

Process & People

white and blue sanitary napkin

Eric Horn

Managing Partner

A few versions back, PTC renamed a cluster property. What had been "master" became "main." Small change, well documented, entirely unrelated to any particular upgrade.

Get it wrong during a version transition and your cluster never comes back online.

On a recent aerospace engagement, our architect knew about that change. He had read it. It simply did not occur to him in the moment — which is what happens to everyone somewhere around hour nine of a maintenance window. The AI agent monitoring the process caught it in seconds, traced the failure back to the renamed property, and said which value needed setting.

That is a good story about AI. But it is a better story about something else.

Long-lived systems accumulate drift

Running that upgrade eight times, in full, turned up a considerable amount of configuration that was simply wrong. Not wrong for 13.1 — wrong, full stop, and quietly wrong for a long time.

Settings left behind by versions that no longer existed. Properties that had been renamed upstream and never updated locally. Configuration that somebody introduced for a reason nobody currently at the company could explain.

None of it was causing a visible outage. All of it was sitting there, and some of it was blocking the upgrade.

This is normal. Any Windchill environment that has been in service for years, through multiple version transitions and multiple administrators, carries this. It is not a sign of a badly run system. It is what happens to systems that are used.

Why it surfaces during an upgrade

Configuration drift is invisible during normal operation, because normal operation exercises a narrow, well-worn path. Everything on that path works, so nothing looks wrong.

An upgrade exercises everything at once. Application, web server, runtime, database, operating system, cluster configuration, vault paths, indexes. Anything that has been quietly wrong for three years is suddenly load-bearing.

Which is why upgrades are so often described as having "gone badly" when what actually happened is that an upgrade discovered pre-existing conditions. The upgrade did not break the environment. It read it out loud.

The audit nobody scheduled

The practical effect of eight rehearsals was that the environment got measurably healthier before the production weekend, in ways that had nothing to do with going from 12 to 13.1.

Under a traditional one-or-two-rehearsal plan, most of that would have stayed hidden. You would have found the items that happened to sit directly in the upgrade path, fixed those, and left the rest for a future engineer to trip over.

Repetition is what converted an upgrade into an audit. Not intelligence — repetition. The agent could read logs, stack traces, database state and directory contents in parallel across every run, and a two-person team monitoring a live upgrade has never had the hours to do that properly.

What to take from this

If you are planning an upgrade, budget for the possibility that you will find things unrelated to it. That is not scope creep. It is a bill that has been accruing quietly, and an upgrade window is a comparatively cheap place to settle it.

And if your Windchill environment has been in service for years without a thorough review, you already have some version of this. The question is only whether you find it on your schedule, or during a maintenance window with the business waiting.

Element implements, upgrades and rescues PTC Windchill for regulated manufacturers. If your environment has drifted and you would rather find out now than on a Saturday night, we can help.

About the author

Eric Horn

Eric Horn is Managing Partner at Element Consulting and a PTC Certified Windchill Implementation Practitioner with 20+ years in PLM across aerospace, industrial, and medical.