Welcome to the M365.FM — your essential podcast for everything Microsoft 365, Azure, and beyond. Join us as we explore the latest developments across Power BI, Power Platform, Microsoft Teams, Viva, Fabric, Purview, Security, and the entire Microsoft ecosystem. Each episode delivers expert insights, real-world use cases, best practices, and interviews with industry leaders to help you stay ahead in the fast-moving world of cloud, collaboration, and data innovation. Whether you're an IT professional, business leader, developer, or data enthusiast, the M365.FM brings the knowledge, trends, and strategies you need to thrive in the modern digital workplace. Tune in, level up, and make the most of everything Microsoft has to offer. M365.FM is part of the M365-Show Network.
Become a supporter of this podcast: https://www.spreaker.com/podcast/m365-fm-modern-work-security-and-productivity-with-microsoft-365--6704921/support.
Advertise on M365.FM - Modern work, security, and productivity with Microsoft 365
Advertise on M365.FM - Modern work, security, and productivity with Microsoft 365 to promote your brand to thousands of podcast listeners
Unlock M365.FM - Modern work, security, and productivity with Microsoft 365 podcast Email contact info, Listeners & Audience details
Email contact information
Direct podcast contact details
Listeners
Audience numbers & engagement insights
Audience details
Podcast Insights
Podcast episodes
Check latest episodes from M365.FM - Modern work, security, and productivity with Microsoft 365 podcast
Why Continuous Improvement Needs Better Data
2026/10/05
Continuous improvement is supposed to create learning that compounds over time. In many factories, however, improvement work still depends heavily on workshops, spreadsheets, isolated reports, and what people remember from the previous shift. A Kaizen event can create visible progress, but a few weeks later the same loss often appears again under slightly different production conditions. The issue is not always the quality of the improvement idea. The bigger problem is that teams often cannot prove whether the countermeasure actually worked, where it worked, and under which conditions it stopped working.
WHY KAIZEN IMPROVEMENTS OFTEN DISAPPEAR
A workshop creates focus for a few days. Teams map the process, identify waste, assign actions, move tools, change checklists, or adjust handoffs. Then normal production pressure returns. The supervisor has another urgent order, maintenance has another fault, planning changes the sequence, and quality puts another batch on hold. The improvement action remains somewhere in an Excel file or project tracker instead of becoming part of the operating rhythm. The important question is therefore not whether an action was completed, but whether the production condition actually improved.
PDCA NEEDS A REAL CHECK STEP
Plan, Do, Check, Act sounds simple, but many organizations effectively run Plan, Do, and Move On. PLAN should define a testable problem, the current condition, the expected improvement, and the hypothesis behind the countermeasure. DO means testing the countermeasure under real production conditions while recording enough context to understand what actually happened. CHECK means comparing the expected result with real production evidence. ACT means standardizing the change when the evidence supports it, or adapting, narrowing, or reversing it when it does not.
• Did the loss actually decrease?
• Did the problem simply move somewhere else?
• Did the change improve availability while damaging quality?
• Did it work across different products, crews, and shifts?
• Did maintenance, material, scheduling, or another process change influence the result?
A single successful production run is not proof. Continuous improvement needs enough evidence to separate a repeatable improvement from a lucky shift.
GEMBA AND DATA NEED EACH OTHER
Data does not replace Gemba. Operators, supervisors, technicians, planners, and quality teams understand production conditions that systems often cannot capture. A machine record may show a ten-minute stop, while an operator knows that the stop happened because a component felt wrong, a normal material route was blocked, or the previous shift left the station in an unusual condition. At the same time, observation alone shows only one shift, one event, or one version of the problem. Connected operational data makes it possible to test whether an observation repeats across orders, products, machines, shifts, material batches, and longer periods of time.
Gemba helps teams identify where to look and which questions matter. Operational data helps test those questions across the real production pattern. Standard work then carries the learning forward so the next shift does not start from zero.
START WITH THE IMPROVEMENT QUESTION
One of the biggest mistakes in manufacturing analytics is beginning with the data that happens to be available. A modern production line can generate huge volumes of information from PLCs, sensors, MES systems, ERP, quality systems, maintenance applications, and spreadsheets. More data does not automatically create better decisions. A useful improvement process starts with the production question the team actually needs to answer.
Instead of asking “What data do we have?”, ask “What production question are we trying to answer?” A statement such as “changeovers take too long” is still too broad. A better question would be: Why does changeover time vary significantly for the same product family on the same production line? That immediately helps define the context that matters.
• Previous product
• Next product
• Production order
• Resource or line
• Tool configuration
• Material
• Shift or crew
• Setup start
• Restart time
• First acceptable unit
• Stable production
• Quality results
• Machine alarms
The goal is not to collect everything. The goal is to collect the smallest set of facts capable of changing the improvement decision.
ERP EXPLAINS THE PLAN
ERP provides the commercial and planning context around production. It can show planned quantities, routing, customer commitments, material requirements, order priority, and the intended production sequence. That context matters because production conditions constantly change. A countermeasure might genuinely reduce setup time while delivery performance still deteriorates because the production mix changed or planners introduced more frequent product transitions.
ERP therefore helps explain what was supposed to run, in what sequence, for which demand, with which routing and materials. But ERP primarily describes intent. It does not necessarily describe what physically happened on the shop floor.
MES EXPLAINS WHAT ACTUALLY RANThe Manufacturing Execution System fills part of that gap. MES can provide the execution history behind an order, including actual operation start and finish, produced quantity, rejects, holds, rework, resources, operator transactions, material consumption, traceability, and reason codes. This allows improvement teams to connect losses with real orders and operations instead of relying on memory or daily averages.
MES data still needs interpretation. Transactions may be entered late, different crews may use reason codes differently, and a timestamp may represent when an operator confirmed something instead of the exact physical moment when it happened. The system provides evidence, but the process gives that evidence meaning.
MACHINE AND IOT DATA EXPLAIN THE PHYSICAL PROCESS
When teams need to understand what happened inside an operation, machine data becomes important. PLC and IoT signals can expose machine states, cycle times, alarms, speed, temperature, pressure, interlocks, motor load, restart behavior, and time to stable production. MES may show that an operation restarted at a particular time, while machine data explains what happened during the minutes before stable output returned.
• Repeated alarms
• Reduced speed after restart
• Temperature recovery
• Manual adjustments
• Interlocks
• Unstable cycle times
Machine data without production context can still be misleading. An alarm becomes much more useful when it can be connected to the work order, product, operation, resource, tool condition, material, and quality result surrounding the event.
THE SHOP FLOOR ADDS MEANING
Operators, maintenance technicians, supervisors, and quality teams fill the gaps that automated systems cannot. A reason code such as “material issue” may describe very different situations in practice. The material may have arrived late, behaved differently, carried an incorrect label, been damaged, or created a downstream problem that only became visible later.
Structured codes help categorize events. Human context helps explain them. Data capture therefore needs to fit the reality of production instead of interrupting it. The useful question is not how much information an operator can enter, but what is the smallest human input that would make the next improvement decision better.
YOU NEED A SHARED FACTORY LANGUAGE
ERP, MES, maintenance systems, historians, and spreadsheets can all describe the same production event differently. One system may call a resource “Line 3,” another may use “LINE03,” maintenance may track several separate assets inside the line, and a historian may still use an old PLC tag. All of these identifiers can technically be correct while still making cross-system analysis unreliable.
The same problem applies to definitions. “Production complete” might mean the machine finished, MES posted the quantity, quality released the batch, the product was packed, or the order was ready to ship. A dashboard cannot decide which definition is correct. The organization has to define that shared meaning. Why Continuous Improvement Need…
THE DATA MODEL BEHIND CONTINUOUS IMPROVEMENT
A useful manufacturing data model preserves meaning as information moves between systems. Products, batches, work orders, resources, process steps, shifts, tools, materials, quality outcomes, maintenance conditions, production events, and timestamps need managed relationships so teams do not rebuild the same logic every time they investigate a problem.
• Product
• Product family
• Work order
• Batch
• Process step
• Resource
• Machine or asset
• Shift
• Tool
• Material
• Quality result
• Maintenance condition
• Production event
• Time
Definitions also need ownership and versioning. Downtime, scrap, rework, good count, target rate, and changeover duration should not quietly mean different things in different reports. Products change, recipes change, tools change, machines are upgraded, routes change, and approved target rates change. Without versioning, today's process rules can accidentally be applied to historical production data.
Become a supporter of this podcast: https://www.spreaker.com/podcast/m365-fm-a-microsoft-mvp-podcast-by-mirko-peters--6704921/support.
IoT Hub Message Routing vs Event Grid — Why Telemetry and Events Are Not the Same Problem
2026/10/02
A temperature reading, a vibration sample, a machine cycle, and a device disconnect may all be described as events, but they do not represent the same architectural problem. In industrial IoT and manufacturing environments, confusing continuous telemetry with discrete events can create noisy workflows, incomplete production histories, unnecessary processing, and systems that react without enough context.
In this episode, we break down the architectural difference between Azure IoT Hub Message Routing and Azure Event Grid by following a realistic manufacturing scenario. A press line continuously sends temperature, vibration, cycle count, energy consumption, and machine-state information through an industrial gateway. Those measurements create an operational history that engineers, data teams, maintenance teams, and production systems may need to analyze later. A device disconnect is different because it represents a change that may require another system or person to react.
TELEMETRY IS A RECORD, NOT AN ALERT
Telemetry represents repeated measurements over time. A single temperature value or vibration measurement usually tells you very little on its own. The real information exists in the sequence: how quickly values changed, what the machine was doing at that moment, what happened before a stop, whether measurements disappeared during a network interruption, and whether the same behavior appeared in earlier production runs.
Typical industrial telemetry includes:
• Temperature, vibration, pressure, energy consumption, and current draw
• Machine states such as running, idle, stopped, or faulted
• Cycle counts, production counters, and process measurements
• Source timestamps, device identifiers, sequence numbers, and correlation information
Those records may later support condition monitoring, quality investigations, energy analysis, OEE calculations, Microsoft Fabric analytics, Power BI reporting, and production optimization. That is why telemetry needs retention, replay, duplicate handling, independent consumers, and a reliable way to reconstruct the production timeline.
IOT HUB MESSAGE ROUTING AS THE TELEMETRY DATA PLANE
Azure IoT Hub provides the controlled device-to-cloud boundary. Devices and gateways authenticate with their own identities, send device-to-cloud messages, maintain device-management state, and can participate in controlled cloud-to-device communication.
Once a telemetry message reaches IoT Hub, Message Routing determines where that data should go. Routing can inspect message properties, system properties, parts of the message body, and device twin information. This makes it possible to separate production telemetry, energy measurements, diagnostics, or other message classes before they reach downstream consumers.
A common architecture might look like this: Industrial gateway → Azure IoT Hub → Message Routing → Event Hubs or Storage → Processing → Microsoft Fabric.
One consumer may perform near-real-time analysis while another keeps a raw archive. A third consumer may prepare curated operational data for Microsoft Fabric. Each consumer can work independently without turning every telemetry reading into a workflow invocation.
WHY ORDERING MATTERS
Industrial telemetry is particularly sensitive to sequence. Imagine a machine reporting that it entered a running state, then transmitting several cycle counts, followed by a process deviation and finally a stopped state. If those records are reconstructed incorrectly, a downstream system could conclude that the machine produced parts while stopped or that a process deviation happened after production had already ended.
The same problem affects downtime calculations, OEE, production counts, and condition monitoring. You therefore need to distinguish between when the source observed something, when IoT Hub received the message, and when a downstream system processed it.
A robust telemetry architecture should therefore consider:
• Source timestamps and cloud receipt timestamps
• Stable partitioning appropriate to the asset or workload
• Message IDs or sequence numbers for duplicate detection
• Idempotent consumers capable of handling at-least-once delivery
Network interruptions make this especially important. A gateway may buffer telemetry and send it when connectivity returns, which means Azure arrival time may be much later than the actual machine timestamp.
EVENT GRID SOLVES A DIFFERENT PROBLEM
Azure Event Grid is designed around publish-and-subscribe notifications. Instead of continuously reconstructing the state of a machine from thousands of readings, Event Grid tells interested systems that something changed and gives subscribers an opportunity to react.
Examples in an IoT environment include device creation, device deletion, connection, disconnection, or carefully selected telemetry-derived conditions. A device-created event might trigger an asset onboarding process. A device-disconnected event might start a technical investigation. A detected engineering condition might trigger a maintenance workflow.
Typical subscribers include:
• Azure Functions
• Logic Apps
• Webhooks
• Security workflows
• Asset-management systems
• Operational applications
The fundamental architectural difference is simple: telemetry asks what an asset has been doing, while an event asks what changed and whether something should react.
EVENTS SHOULD START INVESTIGATIONS, NOT DEFINE REALITY
Event Grid should generally be treated as a notification mechanism, not as the final authoritative state of the physical world. Event delivery can be repeated, and events are not something a subscriber should blindly interpret as the final current state without checking.
A device-disconnected event should therefore usually trigger a state check. An Azure Function might verify the current device state, inspect recent telemetry, check relevant registry information, and then decide whether the investigation should remain open.
Handlers should be designed around:
• Idempotent processing
• Current-state verification
• Event identity and timestamps
• Clearly defined ownership of the resulting action
This matters because duplicate events should not create duplicate tickets, duplicate alerts, or conflicting operational records.
A DEVICE DISCONNECT DOES NOT MEAN THE MACHINE STOPPED
One of the most important distinctions in industrial IoT is the difference between cloud connectivity, data-collection health, and production state. If an industrial gateway disconnects from IoT Hub, the cloud has lost visibility into that gateway. That does not automatically mean that the physical machine stopped.
The PLC may continue controlling the equipment locally while the gateway temporarily loses connectivity. Production may continue normally, and the gateway may even buffer telemetry and upload it later. The opposite can also happen: a gateway may remain connected to Azure while its connection to the PLC or local equipment has failed.
A useful architecture therefore separates cloud connection state, gateway health, machine production state, and MES work-order state. Only by combining those sources can the business determine whether a technical connectivity issue actually affected production.
FILTERING IN IOT HUB VS FILTERING IN EVENT GRID
Both technologies support filtering, but the purpose is different. IoT Hub Message Routing filtering answers which operational messages should enter a specific data path. Event Grid subscription filtering answers which notifications a particular subscriber should receive.
For example, production telemetry may go to one Event Hubs stream while energy data goes to another destination. At the same time, a support workflow may subscribe only to disconnected events from a particular group of gateways.
Confusing those two types of filtering often results in an architecture where every subscriber creates its own interpretation of the telemetry stream. A cleaner design keeps broad telemetry classification governed centrally and keeps Event Grid subscriptions focused on clearly defined reactions.
THE MANUFACTURING CONTEXT LIVES OUTSIDE THE MESSAGE BUS
Neither IoT Hub nor Event Grid knows what a machine means to the business. A device ID may identify a gateway, but it does not inherently know which production line the gateway belongs to, which machine it observes, which work order is running, whether a stop is planned, or whether production is already at risk.
That context usually comes from multiple systems. MES provides execution context such as work orders, operations, recipes, production state, and shift information. ERP provides demand, commitments, inventory, and planning information. An asset model connects devices, gateways, PLCs, machines, cells, lines, and sites.
In more advanced architectures, a digital twin or knowledge graph can help maintain those relationships. The result is a more reliable operational picture in which telemetry explains what the equipment reported, MES explains what production was doing, ERP explains why it matters, and the asset model explains how the technical components relate to the physical plant.
Become a supporter of this podcast: https://www.spreaker.com/podcast/m365-fm-a-microsoft-mvp-podcast-by-mirko-peters--6704921/support.
Building Azure That Survives the Real World — with Mike Martin [MVP]
2026/10/01
Azure architecture looks easy when everything works.The real test starts when dependencies fail, regions become unavailable, traffic spikes, assumptions turn out to be wrong, requirements change, and someone eventually asks the uncomfortable question: why did we design it this way in the first place?In this episode of M365.FM, Mirko Peters talks with Mike Martin [MVP] about what Azure architecture looks like when it has to survive real production conditions rather than just look good on a diagram.Mike brings decades of experience across development, infrastructure, architecture, leadership, coaching, training, and Microsoft Azure. One of his strongest observations is that many of the problems architects face today are not actually new. DNS still breaks. IP dependencies still matter. Costs still become a problem. Integrations still fail. Dependencies still disappear at the worst possible moment.What has changed is the level of complexity we build around those problems.
FROM VISUAL BASIC TO MODERN AZURE ARCHITECTURE
Mike looks back at nearly three decades in IT, starting as a Visual Basic developer in the 1990s and moving through distributed systems, networking, enterprise software, infrastructure, and eventually Azure.His key observation is simple: the industry keeps solving many of the same fundamental problems, but the architectures around them have become much more complex.Modern systems have moved from client-server applications to distributed architectures, cloud platforms, containers, microservices, Kubernetes, hybrid environments, and now AI-assisted development.That creates enormous possibilities, but it also creates a new risk: overengineering.Mike argues that many teams today make simple problems unnecessarily complicated. Good architecture is often not about adding more technology. It is about knowing what not to add.
WHAT DOES AN AZURE ARCHITECT ACTUALLY DO?
For Mike, architecture is not about choosing the largest number of Azure services or producing an impressive diagram.It is about understanding which components belong together, which ones should be avoided, which ones are necessary, and how to build something that remains maintainable, scalable, secure, and resilient.Architecture includes much more than compute.
IdentityNetworkingData flowsSecurityMonitoringIntegrationDependenciesScalabilityOperationsRecoveryDeploymentCostMike also challenges the idea that cloud-native automatically means Kubernetes or containers.Azure provides many managed and native services that can solve problems without introducing unnecessary operational overhead.The architect’s role is to understand the complete solution and choose the simplest architecture that still satisfies the real requirements.
START WITH BUSINESS REQUIREMENTS, NOT AZURE SERVICES
One of the most important lessons in this episode is simple: do not start with technology.Before deciding between Azure Kubernetes Service, App Service, Azure Functions, containers, Service Bus, or another platform, teams should first understand what the solution actually needs to do.Questions to ask:
Is it internal or customer-facing?How many users will depend on it?Does it need to scale globally?How long can it be unavailable?How much data can the business afford to lose?Which compliance requirements apply?What happens if the application disappears for several hours?Who is affected?What level of operational support is required?These questions lead directly into concepts such as SLAs, SLOs, RTOs, and RPOs.They also determine whether the architecture should be single-region, multi-region, active-active, active-passive, or something much simpler.
RTO AND RPO WITHOUT THE BUZZWORDS
RTO and RPO are often discussed as technical acronyms, but their real meaning is business-oriented.
RTO — Recovery Time Objective: How quickly must the system return after a failure?RPO — Recovery Point Objective: How much data loss is acceptable?A system used for non-critical monitoring may tolerate several hours of downtime or lost data.A system supporting first responders, financial operations, commerce, or critical infrastructure may require recovery in minutes.The important point is that these numbers should not be invented by the architect.They should come from the actual business impact of failure.Once those requirements are understood, they can be translated into technical design decisions.
RESILIENCE IS NOT THE SAME AS HIGH AVAILABILITYMike uses a simple analogy to explain the difference between availability and resilience.A highly available system may have another component ready to take over when the primary one fails.A resilient system is designed to absorb problems, continue functioning, recover gracefully, and return to normal without collapsing completely.That distinction matters.An application can technically be available while still providing a poor experience.It may be:
SlowThrottledPartially unavailableDependent on a failing backendAffected by an integration issueUnder regional pressureReal resilience therefore requires more than uptime.It requires an architecture that can cope with load, application errors, partial failures, regional issues, broken dependencies, operational incidents, and recovery after the incident has passed.
AZURE DOES NOT MAKE YOUR APPLICATION RESILIENT AUTOMATICALLY
One dangerous assumption is that because Azure itself is highly available, any application running on Azure automatically inherits that resilience.It does not.Microsoft provides the services and capabilities that make resilient architectures possible.Customers still need to design for resilience.This is where architectural patterns matter:
Circuit breakersLoose couplingAsynchronous processingIndependent scalingMultiple availability zonesMultiple regionsFailover mechanismsMonitoringInfrastructure as CodeTested recovery proceduresMicrosoft provides the building blocks.The architecture determines whether those building blocks actually create a resilient system.
WHY ASYNCHRONOUS DESIGN MATTERS
One of the strongest examples in the episode comes from systems that experience predictable traffic spikes, such as government tax portals.A common design problem occurs when every action depends on a synchronous backend operation.The user clicks a button.The frontend waits for the backend.The backend waits for another service.Another dependency slows down.Eventually the entire user experience is affected.Mike explains why loosely coupled systems are often more resilient.Instead of forcing the frontend to wait for every backend operation, applications can place work into queues or event-driven systems and allow components to process tasks independently.That creates several advantages:
Different components can scale independentlyOne service can fail without taking down the whole platformBackends can process work asynchronouslyTraffic spikes become easier to absorbUser-facing components become less dependent on backend timingFailures become easier to isolateTHE MYTH OF 100 PERCENT UPTIME
Another major topic in the discussion is the obsession with 100 percent availability.Mike is very clear: 100 percent uptime is not a realistic architecture target.A solution is usually composed of multiple services.Each service has its own SLA.Once several dependencies are combined, the effective availability of the complete system changes.A better architecture focuses on:
Fallback scenariosRecovery proceduresRedundancyFailoverMonitoringMitigation strategiesTested recovery plansThe better question is not: how do we guarantee zero downtime?The better question is: what happens when something fails, and how quickly can the business continue operating
WHEN ANOTHER NINE BECOMES TOO EXPENSIVE
More availability is not automatically better.Every additional level of resilience introduces cost, infrastructure, monitoring, operational complexity, and management overhead.Whether that investment makes sense depends entirely on the business impact of failure.An internal holiday request system can probably tolerate several hours of downtime.An order-processing platform handling millions in transactions cannot.Questions that matter:
How much money is lost during downtime?How many users are affected?What happens to customer trust?What happens if an API becomes slow?What happens if orders cannot be processed?What does another level of redundancy actually cost?Is the additional availability worth that cost?Reliability is ultimately an economic decision as much as a technical one.
Become a supporter of this podcast: https://www.spreaker.com/podcast/m365-fm-a-microsoft-mvp-podcast-by-mirko-peters--6704921/support.
Why Your Production Schedule Is Wrong Before It Even Starts
2026/09/28
Your production schedule can look perfectly reasonable when it leaves the ERP system. Order dates line up, routing times make sense, capacity appears available, and material status suggests that production is ready to go. The problem is that the schedule is still only an assumption about the future.The moment production begins, the factory starts generating new facts. Material arrives later than expected, a batch is still waiting for quality inspection, a fixture is unavailable, an operator with a required qualification is missing, or a machine loses capacity because of a short interruption. The original schedule may have been correct when it was created, but the conditions behind it can change within minutes.This episode explores why static production schedules lose accuracy so quickly, why ERP planning is not necessarily the problem, and why modern manufacturing needs a closed feedback loop connecting planning, shop-floor execution, production data, learning, and replanning. The central argument is that a schedule is a decision about what production should try to do based on the information available at that moment. It is not a guaranteed description of what the factory will actually be able to execute. Why Your Production Schedule Is…
WHY PRODUCTION SCHEDULING BREAKS DOWN
Every scheduled production operation contains multiple hidden assumptions. A planned start time assumes that the previous job finishes on time, the machine remains available, the required material is usable, the correct tool or fixture is ready, and a qualified operator is present.It may also assume that setup duration remains realistic, that actual cycle time stays close to the routing standard, that quality releases the material as expected, and that another more urgent order does not suddenly compete for the same resource.That means a production schedule is not simply a table containing orders, dates, quantities, and machines. It is a collection of assumptions about future operating conditions.Typical assumptions include:
Machine availabilityMaterial readinessTool and fixture availabilityOperator qualificationsSetup durationCycle timeQuality releaseResource capacityProduction sequenceCustomer prioritiesThe schedule becomes unreliable when those conditions change but the decision is not updated.
ERP PLANNING IS NOT THE PROBLEM
ERP remains one of the most important systems in manufacturing. It connects customer orders, inventory, bills of material, purchasing, routings, work centers, due dates, and business commitments.ERP provides the commercial intent behind production. It tells the organization what should be produced, which demand needs to be covered, which materials are required, and which customer commitments matter.The limitation appears when ERP planning is expected to understand every operational condition on the shop floor at every moment.A work center may appear available while the required fixture is still installed somewhere else. A material receipt may exist in ERP while the batch is still waiting for inspection. A person may appear on the workforce calendar while lacking the specific qualification required for the next operation.The ERP system is not necessarily wrong. The factory has simply produced newer information.
PRODUCTION PLAN VS DETAILED SCHEDULE VS DISPATCH LIST
One reason production planning becomes confusing is that several different decisions are often described using the same word: schedule.A production plan usually works at a broader level. It determines what demand needs to be covered, which product families should be produced, and whether enough capacity and material appear to exist across a longer planning horizon.A detailed production schedule moves closer to execution. It determines which operation should run on which resource, in what sequence, and within which time window.A dispatch decision operates even closer to the shop floor. It answers the practical question: what should this machine, operator, or work center run next based on the conditions we know right now?These decisions are connected, but they are not identical. The closer production gets to execution, the more important current operational conditions become.
WHY EXCEL BECOMES THE UNOFFICIAL MANUFACTURING CONTROL SYSTEM
When the official production schedule no longer matches the factory, planners frequently move into Excel. The reason is simple: Excel reacts faster.A planner can change priorities, reorder jobs, add comments, highlight material issues, record tooling problems, and send a revised sequence within minutes.That flexibility is valuable when the factory needs an operational decision immediately.The spreadsheet itself is therefore not necessarily the underlying problem. It often exposes a capability that the formal production system does not currently provide.The deeper issue appears when the production decision becomes fragmented across different places:
ERP contains the original production intent.Excel contains exceptions and manual schedule changes.Supervisors hold the immediate operational sequence.Operators hold practical knowledge about what can actually run.Emails, calls, whiteboards, and shift handovers carry additional context.Once this happens, no single system represents the complete production state.The episode describes Excel as the place where exceptions, calls, and practical decisions accumulate because it can respond faster than the formal process. Why Your Production Schedule Is…
THE CLOSED-LOOP PRODUCTION MODEL
The solution is not simply to regenerate the schedule more often. Manufacturing needs a controlled feedback loop.A practical closed-loop production model follows five stages:
PLAN — Use demand, capacity, routings, material, and business priorities to create the initial production decision.EXECUTE — Release work to the shop floor and observe what actually happens.MEASURE — Capture events that materially change the assumptions behind the plan.LEARN — Compare planned assumptions with repeated production behavior.REPLAN — Use the current production state to create the next feasible dispatch decision.This changes the objective of scheduling. Instead of trying to create one perfect schedule and defend it against reality, the organization builds a process capable of reacting when reality changes.
EXECUTION PRODUCES FACTS THAT PLANNING COULD NOT KNOW
Once production begins, every operation creates information that can change the remaining schedule.An operation may start late. A setup may take longer than expected. A machine may stop for twenty minutes. A quality issue may block a batch. Scrap may reduce the quantity available for the next operation. Rework may create additional demand on an already constrained resource.These are not just historical records for a weekly production report.They change what the factory can do next.For example, if an operation finishes one hour late, every downstream operation may shift. If a batch is blocked by quality, another work center may suddenly have unused capacity. If rework sends material back through an earlier operation, the rework now competes with planned production for the same machine.A useful scheduling system therefore needs execution feedback while the schedule is still active.
MEASURE AT THE LEVEL OF THE PRODUCTION DECISION
Manufacturing environments already generate huge volumes of machine and process data. The problem is not always a lack of data.The problem is often a lack of context.A machine reporting a twenty-minute stop tells you that capacity was lost. It does not automatically tell you which customer order was affected or which production decision should change.To support production scheduling, an event should ideally connect to information such as:
Production orderOperationProductBatchResourceShiftEvent timeReason codeCurrent queueDownstream dependencyThe episode emphasizes that production measurement should operate at the same level where scheduling decisions happen: order, operation, resource, and time. Why Your Production Schedule Is…Department-level totals may tell you that output was below plan, but they cannot necessarily explain which operation consumed unexpected capacity or whether the next scheduled job is still feasible.
OEE DOES NOT DECIDE WHAT SHOULD RUN NEXT
Overall Equipment Effectiveness remains a useful manufacturing metric because it helps organizations understand equipment availability, performance, and quality.But OEE and production scheduling solve different problems.A machine can have strong OEE while production still misses an important customer delivery because the wrong work was placed in front of the bottleneck.Likewise, poor OEE does not automatically tell the planner what sequence should run for the rest of the shift.Production scheduling needs additional information such as actual cycle time, queue time, material readiness, tool availability, workforce qualifications, quality status, setup requirements, and customer priorities.OEE can help explain resource performance. Scheduling must determine how limited resources should be used across competing work.
Become a supporter of this podcast: https://www.spreaker.com/podcast/m365-fm-a-microsoft-mvp-podcast-by-mirko-peters--6704921/support.
Microsoft 365 Copilot Without the Hype: Adoption, Governance & Getting Your Tenant Ready with Paul Keijzers [MVP]
2026/09/24
Microsoft 365 Copilot can help people find information and get work done faster, but its answers depend on the content and permissions already in an organization’s Microsoft 365 tenant. In this episode of the M365 FM podcast, Mirko Peters speaks with Microsoft MVP Paul Keijzers, founder of KB Works, about preparing Microsoft 365 for Copilot in a practical, responsible way. They look beyond demos and licensing to the foundations that shape Copilot results: SharePoint, Microsoft Teams, OneDrive, data quality, access controls and governance. Copilot can make existing access issues more visible. Files that have been shared broadly, outdated documents, old versions and content with unclear ownership may all affect what employees can find. Paul explains why organizations should review their sharing settings and understand where sensitive or unnecessary access exists before rolling Copilot out widely. SharePoint oversharing reports can help identify potential issues across SharePoint, Teams and OneDrive, although organizations still need to check whether flagged sharing is appropriate for their business.
GOVERNANCE, LEGACY DATA AND INFORMATION PROTECTION
The conversation explores how to manage years of legacy SharePoint content. Old project files may need to be retained for legal or business reasons, but that does not always mean they should appear in everyday search or Copilot results. Organizations can consider archiving content or restricting access, depending on how often it needs to be used and how the tenant is structured. Paul also highlights the need to plan for Microsoft 365 backup and recovery, including how a company can retrieve its data if it changes backup providers. Microsoft Purview is part of the wider governance picture. Paul recommends reviewing sensitivity labels, checking whether they are applied consistently, and considering automatic labeling where appropriate. Data Loss Prevention policies can help protect personal and sensitive information, while Power Automate DLP policies need attention when Copilot uses connectors to work with services such as Jira. These controls should fit the organization’s actual needs, since a global company and a small business may require different policies.
SHAREPOINT STRUCTURE AND COPILOT ADOPTION
Good information architecture remains important even as AI becomes better at understanding natural language. Paul discusses using metadata and content types to make information easier to organize and retrieve, while keeping SharePoint libraries practical for employees. His guideline is to avoid overly deep folder structures and excessive metadata fields, since people are less likely to maintain a system that is too complicated. Copilot can help suggest or populate information, but organizations still need a clear structure and reliable content. Copilot adoption also requires ongoing support. Rather than delivering a single training session and expecting employees to figure out the rest, organizations can share regular tips, demonstrate useful prompts and agents, and help teams solve real daily frustrations. Paul recommends starting with the work people actually do, then identifying where Copilot could save time or improve an outcome. Adoption plans should be adapted to the size and working practices of each organization instead of copied wholesale from a framework designed for a different environment.
AI VALUE, AGENTS AND A PRACTICAL PATH FORWARD
Mirko and Paul also discuss how organizations can assess the value and cost of AI, how agents may increasingly work alongside employees, and why administrators need ways to discover and manage agents in their environments. Paul shares his Focus Week concept in Portugal, where teams spend dedicated time working on Microsoft 365, SharePoint, governance or Copilot challenges away from their usual workplace interruptions. The central message for IT leaders is that Copilot readiness starts with understanding the Microsoft 365 tenant: who can access what, how information is organized, which content should be retained or surfaced, and how employees will learn to use AI in their work. Review sharing settings, improve information governance and connect adoption to real business needs before treating Copilot as simply another license to deploy.
Become a supporter of this podcast: https://www.spreaker.com/podcast/m365-fm-a-microsoft-mvp-podcast-by-mirko-peters--6704921/support.
When Microsoft Messaging Looks Secure but Isn’t — Exchange Server, Hybrid and Microsoft 365 Security with Thomas Stensitzki [MVP]
2026/09/24
Microsoft Exchange and Microsoft 365 make it possible to run powerful messaging environments, but moving email to the cloud does not automatically make it secure. In this episode of the M365 Show, host Mirko Peters [MVP] talks with Thomas Stensitzki [MVP] about the security gaps that can hide in Exchange Server, Exchange Online, and hybrid deployments. Drawing on more than 25 years of messaging experience, Thomas explains how identity protection, careful configuration, secure mail flow, and operational discipline work together to protect an organization’s email.
FROM EXCHANGE SERVER TO HYBRID AND EXCHANGE ONLINE
Thomas shares how he built his career around Exchange and why email remains essential to business. He explains how a hybrid setup connects on-premises Exchange Server with Exchange Online, and why that connection needs careful planning across messaging, networking, and security teams. Microsoft 365 provides a working service with default settings, but organizations still need to configure protections such as anti-spam, anti-malware, and anti-phishing to fit their needs.
IDENTITY, ADMINISTRATOR ACCESS, AND BREAK-GLASS ACCOUNTS
Identity security comes first, including for service accounts and other non-human identities. Thomas discusses sensitive Exchange administrator roles, privileged access management, and why administrators should avoid using highly privileged accounts for everyday work. He also explains how to protect emergency or “break-glass” accounts, including the role of FIDO security keys and the need to plan how administrators can regain access during an outage.
LEGACY SMTP, PHISHING, AND COMPROMISED ACCOUNTS
Older applications and devices may still depend on basic authentication or legacy SMTP. Thomas recommends avoiding those methods where possible and describes how an on-premises relay can help route messages from systems that cannot use modern authentication. The conversation also follows a potential attack path from a malicious email to stolen credentials and unauthorized access, highlighting the value of email filtering, separate administrative accounts, and monitoring sign-in activity.
MAIL FLOW, SPF, DKIM, AND DMARC
Understanding the full route an email takes is essential, especially in complex environments that combine gateways, Exchange Server, Exchange Online Protection, and Microsoft Defender. Thomas explains the roles of SPF, DKIM, and DMARC in authenticating messages sent from an organization’s domain. He also recommends using dedicated subdomains for third-party services such as marketing platforms and CRM systems, and using DMARC reports to identify legitimate and suspicious senders.
MICROSOFT DEFENDER AND SECURITY MONITORING
Thomas discusses Microsoft Defender for Office 365 Safe Links and how link protection can help assess a URL when a user clicks it. He also covers Entra sign-in logs, suspicious sign-ins, and impossible-travel alerts. Security tools and scores can guide decisions, but administrators still need to understand what they measure, what their licenses include, and which protections their organization actually needs.
CONFIGURATION, GOVERNANCE, AND OPERATIONAL DISCIPLINE
The discussion moves beyond individual security settings to configuration management, least-privilege access for programmatic tools, and the risks of exposing Microsoft 365 content through APIs or agents. Thomas explains how configuration exports can help teams track changes. He also discusses Microsoft Purview sensitivity labels and data loss prevention (DLP), recommending that organizations plan their rollout carefully because these controls can be difficult to change once they are in production.
BUSINESS CONTINUITY, AI, AND KEEPING EXCHANGE SECURE
Backups alone may not be enough if an organization loses access to its Microsoft 365 tenant, domains, or configuration. Thomas stresses the importance of planning for business continuity before an incident occurs. He also considers how AI may help both defenders and attackers, and describes his consulting work helping organizations update Exchange Server environments and move to Exchange Online or hybrid configurations.
RAPID-FIRE QUESTIONS AND FINAL SECURITY ADVICE
In the rapid-fire round, Thomas chooses Exchange Server, a long-term hybrid architecture, and PowerShell. He also discusses Exchange Server certificate management and recommends Manfred Huber as a future guest. His closing advice is straightforward: keep Exchange environments up to date, follow the Exchange Product Group blog, apply patches, watch for default configuration changes, and prepare for the deprecation of Exchange Web Services in Exchange Online.
Become a supporter of this podcast: https://www.spreaker.com/podcast/m365-fm-a-microsoft-mvp-podcast-by-mirko-peters--6704921/support.
Content Understanding & Document AI - Simply Explained
2026/09/18
Invoices, contracts, receipts, forms, scanned PDFs, and email attachments contain valuable business information — but most automation still struggles to turn those documents into reliable, structured data.
In this episode of M365 FM – Simply Explained, we break down Microsoft Content Understanding, Document AI, OCR, AI Builder, Power Automate, Dataverse, and Power Apps and explain how they work together to transform documents into usable business data and automated processes.
You’ll learn why traditional OCR is only the beginning. OCR can recognize text such as an invoice number or amount, but it does not automatically understand whether a number represents an invoice ID, purchase order, bank account, tax value, or phone number. Document AI adds context, structure, and business meaning to extracted information.
We explain how Microsoft Content Understanding can take documents and images, extract defined fields, classify information, and return structured results that downstream systems can use. Instead of asking AI to summarize an entire document, organizations can define a schema containing fields such as supplier name, invoice date, invoice number, invoice total, document type, and line items.
The episode also shows how Microsoft Power Platform turns document extraction into a complete business workflow. Power Automate can detect new documents in email or SharePoint, send them for extraction, validate the returned information, create records, trigger approvals, and route exceptions to the right person.
Dataverse can store structured document records and process history, while Power Apps can provide a human review interface for correcting uncertain or missing information.
We also cover one of the most important parts of Document AI: confidence scores and human-in-the-loop review. A high confidence score does not automatically mean a value should be trusted.
Critical information such as invoice totals, payment instructions, bank details, or contract dates may still require additional validation against business rules and existing systems.
You’ll discover how validation can check whether suppliers exist, purchase orders match, totals make sense, dates are valid, and duplicate invoices have already been processed.
This combination of AI extraction, validation rules, governance, and human review is what turns Document AI from an impressive demo into a reliable business process.
We also look at practical first use cases including invoice processing, employee onboarding forms, claims, contract expiry dates, supplier documents, procurement workflows, HR documents, legal documents, service requests, and customer forms.
The key is to start with one document type, one clear decision, a defined owner, and a review path for exceptions.
By the end of this episode, you’ll understand the complete Document AI pattern:
Document → Content Understanding → Structured Data → Validation → Power Automate → Dataverse → Human Review → Business Action
The goal is not simply to process more PDFs. It is to stop people searching through documents for basic information and instead move structured, validated data directly into the business process where decisions happen.
Subscribe to M365 FM for practical episodes about Microsoft Content Understanding, Power Platform, Power Automate, AI Builder, Microsoft AI, automation, Copilot, document processing, and the future of work.
Become a supporter of this podcast: https://www.spreaker.com/podcast/m365-fm-a-microsoft-mvp-podcast-by-mirko-peters--6704921/support.
IoT Hub Message Routing vs Event Grid — Why Telemetry and Events Are Not the Same Problem
2026/09/18
A machine sends temperature readings, vibration data, cycle counts, power consumption, and operating states every few seconds. Then the gateway suddenly disconnects. Are all of those messages simply “events”? Technically, you could describe them that way. Architecturally, that can create serious problems. Azure IoT Hub Message Routing and Azure Event Grid solve different problems. One path is designed around preserving and distributing operational data. The other is designed around notifying systems that something changed and may require a response. Treating them as interchangeable can leave you with expensive workflows processing routine sensor data—or important signals buried inside a telemetry pipeline nobody is actively watching. In this episode of M365 FM, we follow a manufacturing machine through a real Azure IoT architecture and explain where IoT Hub, Event Grid, Event Hubs, Microsoft Fabric, Power BI, MES, ERP, Functions, and Logic Apps actually belong.
WHAT YOU WILL LEARN
In this episode, we explore:
Why machine telemetry and discrete business or lifecycle events require different architecture patternsHow Azure IoT Hub Message Routing works as part of a telemetry data planeWhere Azure Event Grid fits into event-driven and reactive architecturesWhy message ordering matters for manufacturing telemetryWhy Event Grid should not become your primary high-volume telemetry busWhy IoT Hub routing should not be forced into every notification workflowHow IoT Hub and Event Grid can work together in the same architectureHow Event Hubs can support independent stream-processing consumersWhy raw telemetry should often be retained for traceability and later investigationHow Microsoft Fabric and Power BI can consume prepared operational dataWhy MES, ERP, and asset models provide context that device data alone cannot provideHow device disconnect events should be interpreted without automatically assuming production stoppedHow duplicate delivery, retries, timestamps, and idempotency affect reliable industrial architecturesHow to design condition monitoring, predictive maintenance, quality traceability, and production-disruption workflowsHow to decide whether a message belongs on the data plane, the response path, or bothTELEMETRY IS A RECORD OVER TIME
Telemetry is not valuable because one temperature reading arrived. It becomes valuable because thousands of readings together describe what happened. A production machine may continuously report:
Temperature and vibration measurementsMotor current and energy consumptionCycle counts and production countersRunning, idle, stopped, or faulted statesSource timestamps and sequence informationDiagnostic and equipment-health informationA single temperature value might mean very little. The sequence around that reading tells the story. Was the machine warming up? Was it already producing? Was vibration increasing at the same time? Did cycle time begin to increase? Did the machine stop shortly afterward? Telemetry therefore needs a path designed around sequence, retention, replay, independent consumers, and traceability.
EVENTS EXIST TO START A RESPONSE
An event serves another purpose. An event says: Something changed. A system or person may need to react. Examples include:
A new device was registeredA gateway disconnected from IoT HubA device reconnectedA device was deletedA monitoring process detected a condition requiring investigationAn inspection completed and another workflow can beginThe recipient usually does not need hours of telemetry before starting the first step. It needs enough information to identify what happened and determine the appropriate response. That response might involve:
Starting an Azure FunctionTriggering a Logic AppOpening a support investigationUpdating an asset recordChecking the current device stateCalling an external application through a webhookNotifying the team responsible for the affected systemThe event starts the investigation. It does not necessarily contain every fact needed to make the final operational decision.
WHY “EVERYTHING IS AN EVENT” BREAKS DOWN
Sending every sensor measurement into event-triggered workflows can look attractive during a proof of concept. Then production scale arrives. Every reading triggers another Function. Another Logic App evaluates something. Another integration receives another message. Maintenance creates its own subscription. Quality creates another. Energy management creates another. Soon, every team has slightly different filtering, state management, retry handling, and storage logic. A temperature measurement is not automatically an incident. It may contribute to an incident later, but the continuous measurements should remain available as evidence. When routine telemetry starts generating constant notifications, users can also begin ignoring alerts because the system has trained them to expect noise rather than actionable information.
WHAT AZURE IOT HUB ACTUALLY DOES
Azure IoT Hub provides the device-facing cloud boundary. It supports areas such as:
Secure device identitiesDevice-to-cloud messagingCloud-to-device communicationDevice managementDevice twinsControlled access between connected equipment and Azure servicesIoT Hub knows that an authenticated device sent a message. It does not automatically understand what that message means to production. A gateway might report: state = running But “running” could mean:
Producing approved partsDry cyclingRunning setupPerforming reworkMoving without materialThat context usually comes from additional systems such as the MES, ERP, asset model, or production application.
Become a supporter of this podcast: https://www.spreaker.com/podcast/m365-fm-a-microsoft-mvp-podcast-by-mirko-peters--6704921/support.
The 5 Pillars of Data Transformation - Simply Explained
2026/09/18
AI was supposed to clear the backlog, accelerate decisions, and give every team a smarter way to work. Instead, many organizations now have Microsoft Copilot, Power BI, Microsoft Fabric, AI agents, and more data than ever before—while important decisions still crawl through meetings because nobody fully trusts the numbers or knows who can act on them. The technology spend keeps rising. The action does not. The problem is often not a lack of AI. It is the absence of an operating model connecting data, meaning, governance, technology, people, and accountability. In this episode of M365 FM – Simply Explained, we break down the five pillars organizations need to build a reliable foundation for data transformation and AI.
WHAT YOU WILL LEARN
In this episode, we explore:
Why AI cannot compensate for unreliable dataHow data governance creates trust before automation beginsWhy data quality should depend on the decision being madeHow Microsoft Purview can support governance and data discoveryHow Microsoft Fabric supports modern analytics and data platformsWhy semantic models matter for Power BI and AIHow conflicting definitions create conflicting dashboardsWhy business glossaries matter for humans and AI agentsHow data ownership affects AI readinessWhy access, security, and permissions must be defined before AI scalesHow Copilot and AI agents depend on trusted business contextWhy human accountability remains critical even when AI generates the answerPILLAR 1: DATA GOVERNANCE – TRUST BEFORE AUTOMATION
Data governance often sounds like policies, compliance meetings, documentation, and bureaucracy. In practice, governance answers a few very simple questions:
Who owns this data?Who is allowed to access it?Where did the data come from?Can we trust it for this particular use case?What are people allowed to do with it?What are AI systems allowed to do with it?Without clear answers, AI does not solve a data problem. It can spread the problem faster. Imagine a leadership team preparing a sales forecast. Sales presents one revenue number. Finance presents another. Both numbers come from systems that appear authoritative. The meeting suddenly stops being about future decisions. Instead, everyone starts arguing about which spreadsheet or dashboard is correct. The underlying problem may be that:
The CRM contains one version of revenueThe finance system contains anotherManual exports introduce additional differencesNobody owns the definition of revenueNobody owns the quality of the source dataNobody can clearly explain which number should drive the forecastThe company ends up debating the past instead of deciding the future.
WHAT HAPPENS WHEN AI ENTERS THE PICTURE?
Now imagine someone asks an AI agent: “Which sales region is falling behind?” The answer may arrive within seconds. But it could be based on:
Duplicate customer recordsOutdated account assignmentsMissing opportunitiesIncorrect forecast stagesOld dataIncorrect permissionsInformation the user should not have been able to accessThe answer can sound confident. That does not automatically make it trustworthy. Governance creates the working agreement around the data before automation starts using it. A strong governance model typically establishes:
Named data ownersClear responsibilitiesData classificationsAccess rulesSource-system documentationData quality expectationsAuditabilityPolicies for sensitive informationRules for AI and automationMicrosoft technologies can support this process. Microsoft Purview can help organizations discover, classify, understand, and govern information. Microsoft Fabric can help bring data together, prepare it, analyze it, monitor it, and make it available for reporting and AI scenarios. But technology cannot decide everything. Organizations still need people to decide:
Who owns customer dataWhich definitions are authoritativeWhat data quality is acceptableWho should have accessWhen an AI-generated answer is safe to useWho remains responsible for the final decisionDATA QUALITY MUST MATCH THE DECISION
Many organizations approach data quality as if every field in every system needs to be perfect. That is rarely realistic. Data quality should instead be evaluated against the business decision being made. For a sales forecast, the most important fields might include:
Opportunity stageExpected close dateForecast amountAccount ownerTerritoryProbabilityCustomer statusOther fields may be less important for that specific decision. The better question is therefore not: “How do we clean all of our data?” The better question is: “Which data must we trust for this decision?” That makes the problem smaller, more measurable, and much easier to manage.
Become a supporter of this podcast: https://www.spreaker.com/podcast/m365-fm-a-microsoft-mvp-podcast-by-mirko-peters--6704921/support.
Your Factory Cloud Bill Is Much Higher Than You Think
2026/09/17
Cloud storage may look cheap. Sending factory data to the cloud may look cheap. But the real cloud bill often starts when that data begins moving. In this episode, we break down the hidden costs behind modern Industrial IoT, manufacturing cloud, edge computing, and factory data architectures — from data egress and cross-region replication to NAT gateways, backups, dashboards, and high-frequency sensor data. A single MQTT stream from a factory can quickly become multiple data flows once telemetry is copied into storage, analytics platforms, dashboards, data science environments, disaster recovery systems, and external applications. The machine generated the data once — but your architecture may move it many times.
Why Factory Cloud Costs Grow So Quickly
One of the biggest mistakes in manufacturing IoT architecture is estimating data volume based on the number of machines or connected devices. The better calculation is: Samples × Bytes × Time × Assets A simple machine-state signal may generate very little data. A vibration sensor sampling at 32 kHz is completely different: a single 16-bit channel can generate roughly 5.5 GB of raw data per day before additional protocol and metadata overhead. This episode explores why an edge-first architecture can dramatically change that equation. Instead of continuously uploading every raw measurement, manufacturers can process data close to the machine, retain detailed evidence locally, detect meaningful changes, create aggregates, and send only the information required by cloud consumers.
What You'll LearnWhy cloud egress costs can become more important than storage costsHow MQTT and IoT telemetry can create multiple downstream data flowsWhy device count is a poor way to estimate factory data volumeHow vibration monitoring can generate gigabytes or terabytes of dataWhy cross-region and cross-zone traffic mattersHow NAT gateways and network routing can increase cloud costsWhy replication, backups, exports, and dashboards create additional data movementHow to identify duplicate factory data pipelinesWhen raw manufacturing data should remain at the edgeHow event filtering and aggregation reduce unnecessary cloud trafficWhy edge computing should be a processing layer rather than a miniature cloudHow to design an edge-to-cloud manufacturing architecture around business decisions rather than raw data volumeEdge Computing vs. Sending Everything to the Cloud
The key architectural question isn't:
“Can we send this factory data to the cloud?”
It's:
“What data actually earns the trip?”
High-rate raw signals such as vibration waveforms, diagnostic traces, and vision data can often remain close to the factory. Filtered events and aggregates can move selectively, while production records, quality outcomes, KPIs, and cross-plant analytics are stronger candidates for centralized cloud platforms. The result is not an argument against cloud computing. It is a more deliberate IT/OT architecture in which edge and cloud have different responsibilities.
Topics Covered
Industrial IoT, IIoT, Edge Computing, Cloud Computing, Manufacturing Data, Factory Data, MQTT, OPC UA, Data Egress, Cloud Costs, FinOps, Azure IoT, AWS IoT, Factory Automation, Predictive Maintenance, Vibration Monitoring, Data Architecture, IT/OT Integration, Smart Manufacturing, Industry 4.0, Data Replication, Cloud Networking, Manufacturing Analytics
Who Should Listen?
This episode is for manufacturing IT leaders, OT engineers, cloud architects, IoT architects, data engineers, plant managers, solution architects, and industrial digitalization teams designing or operating connected factory environments.
If your architecture contains a neat arrow labeled “Factory → Cloud,” this episode explains why that arrow deserves a much closer look.
Become a supporter of this podcast: https://www.spreaker.com/podcast/m365-fm-a-microsoft-mvp-podcast-by-mirko-peters--6704921/support.
From Financial Data to AI-Ready Decisions: Power BI, Microsoft Fabric, Semantic Models & AI Agents with Rishi Sapra [MVP]
2026/09/16
What happens when AI agents start making sense of financial and business data — not just displaying it?In this episode of M365.FM, Mirko Peters talks with Rishi Sapra Microsoft MVP about the architecture required to move beyond traditional dashboards and toward AI-powered, context-aware analytics with Microsoft Fabric, Power BI, Copilot Studio, semantic models, ontologies, and data agents.Organizations already have enormous amounts of information spread across ERP systems, Excel workbooks, Power BI reports, Microsoft Fabric, SharePoint, financial systems, budgets, forecasts, and operational applications.Adding AI on top of that data does not automatically mean the AI understands the business.What does “revenue” actually mean? Which definition of margin should an AI agent use? Which KPIs represent the official version of the truth? And how does an agent understand relationships between customers, products, regions, cost centers, contracts, and business processes?The answer increasingly lies in the context layer between raw data and AI.
FROM SELF-SERVICE BI TO SELF-SERVICE AI
Rishi explains his journey from financial modeling and Excel through Power Query and the early days of Power BI to today's Microsoft Fabric and AI ecosystem.Power BI helped bring business intelligence out of centralized IT departments and into the hands of business users.But the platform has also become significantly more sophisticated.Semantic models, DAX, lakehouses, OneLake, Direct Lake, Copilot Studio, data agents, ontologies, MCP-based tools, governance, and AI now create an architecture that can quickly become difficult for individual business users to understand.AI agents could change that relationship.Instead of requiring every business user to become a data engineer, BI developer, and AI engineer, agents could increasingly build and operate parts of the technical architecture while humans provide the business context.
WHY BUSINESS CONTEXT MATTERS FOR AI
One of the central questions of the episode is surprisingly simple:How well documented are your business processes and decisions?Much of an organization's real knowledge does not live in a database. It exists inside Excel formulas, Power BI measures, SharePoint files, business processes, documentation — and people's heads.For AI agents to produce meaningful business insights, organizations need to capture more than data.They need to capture context.Who is asking the question?What decisions does that person need to make?Which KPIs matter?Which business rules apply?What does a specific metric mean in that particular context?This leads to the concept of persona-driven insights: designing analytics around the decisions and questions of specific business users rather than simply exposing more data.
DATA MODELS VS SEMANTIC MODELS VS ONTOLOGIES
The conversation explores three increasingly important concepts in modern Microsoft analytics architecture.A data model structures the underlying data and relationships.A Power BI semantic model adds business logic, measures, calculations, relationships, and security — creating a governed analytical layer and a reliable source for KPIs.But AI often needs more.An ontology can describe business entities and relationships in a way that allows AI to reason about concepts such as customers, products, stores, employees, regions, contracts, revenue, and business processes.Semantic models help answer:“What is the number?”Ontologies and additional context can help AI investigate:“Why did the number change?”Together, these layers provide much stronger grounding for AI agents.
MICROSOFT FABRIC AS THE DATA FOUNDATION FOR AI
Microsoft Fabric plays a central role in this architecture.OneLake, Lakehouses, Delta tables, semantic models, Direct Lake, Fabric Data Agents, and integration with Copilot Studio can create a unified foundation for structured and unstructured organizational data.The episode also explains why Direct Lake matters.Instead of repeatedly importing and refreshing data into traditional Power BI semantic models, Direct Lake allows Power BI to work directly with data stored in Delta format while maintaining analytical performance.This can significantly simplify the path from enterprise data to analytics and AI.
THE FIVE LAYERS OF AN ORGANIZATIONAL BRAIN
Rishi describes an “organizational brain” built around five interconnected layers:Data — trusted enterprise information and source systems.Logic — DAX, SQL, Python, calculations, KPIs, and business rules.Tools — semantic models, APIs, MCP servers, applications, and other capabilities agents can use.Skills — instructions and business processes describing how agents should use those tools and interpret information.Governance — permissions, policies, security, controls, and rules governing what agents are allowed to do.The goal is not simply to give an LLM access to more data.The goal is to give AI a governed environment in which it understands which data, logic, tools, and processes should be used for a particular business question.
AI AGENTS NEED DETERMINISTIC DATA
Generative AI is powerful because it can reason flexibly.Financial reporting cannot rely entirely on flexibility.Revenue, margins, forecasts, costs, and other business metrics often require deterministic calculations and a governed source of truth.The episode explores why the future of enterprise AI may therefore depend on combining two worlds:Deterministic computing for trusted calculations and business logic.Generative AI for reasoning, interpretation, exploration, and natural-language interaction.Semantic models and Microsoft Fabric can provide the deterministic foundation while AI agents provide the flexible reasoning layer.
FROM DASHBOARDS TO PERSONALIZED INTELLIGENCE
Traditional dashboards tell users what happened.A dashboard might show that revenue decreased by seven percent. The user then needs to drill through dimensions, filters, reports, and datasets to understand why.AI agents can potentially perform much of this exploration automatically.But good storytelling still requires context.The most important number is not always the largest number. A business metric may need to be interpreted relative to revenue, budget, previous periods, organizational structure, or other factors.This is where persona-driven analytics becomes particularly important.The CFO, sales leader, and operational manager may all ask about the same KPI while requiring very different explanations and actions.
COPILOT STUDIO AND THE NEXT GENERATION OF AGENTS
The conversation also explores the evolution of Microsoft Copilot Studio from traditional topic- and knowledge-based chatbot experiences toward more capable agentic systems.Modern agents can potentially combine tools, skills, enterprise data, workflows, and reasoning.That additional capability also introduces additional risk.The more autonomy an agent receives, the more important grounding, permissions, governance, evaluation, and clearly defined instructions become.AI agents should not invent financial numbers or arbitrarily choose data sources.They need trusted semantic models, governed data, explicit skills, and clear guardrails.
MAKING GOVERNANCE GREAT AGAIN
Governance becomes even more important in an agentic organization.Instead of treating governance as a document that employees are expected to read, organizations can increasingly encode governance directly into the environments, policies, skills, and instructions used by AI agents.The discussion explores the idea of treating agents more like digital employees.They need defined responsibilities, approved tools, access permissions, business rules, performance expectations, and boundaries.Governance therefore becomes less about documentation and more about automated enforcement.
THE OPERATING MODEL FOR FABRIC AND AI
Finally, the episode examines how organizations can manage this architecture at scale.A Center of Excellence can provide the enablement layer while a hub-and-spoke model combines centralized governance with decentralized innovation.Certified enterprise data, semantic models, logic, governance policies, and reusable skills can live in governed hubs.Business teams can experiment within their own domains and promote successful assets into the governed enterprise layer.The result is neither completely centralized nor completely decentralized.It is a federated model designed to support both control and self-service.
IN THIS EPISODE
We discuss Microsoft Fabric, Power BI semantic models, data modeling, ontologies, OneLake, Lakehouses, Direct Lake, Delta tables, Fabric Data Agents, Microsoft Copilot Studio, AI agents, persona-driven insights, storytelling with data, deterministic computing, enterprise AI governance, Center of Excellence models, hub-and-spoke architectures, self-service BI, self-service AI, and the idea of building an organizational brain for AI.The bigger question is no longer simply:“How do we build better dashboards?”It is:“How do we give AI enough trusted business context to understand our organization — without allowing it to invent its own version of the truth?”
Become a supporter of this podcast: https://www.spreaker.com/podcast/m365-fm-a-microsoft-mvp-podcast-by-mirko-peters--6704921/support.
Why Your ERP Can't Build an Optimal Production Schedule
2026/09/16
Your ERP can calculate production dates, explode demand through MRP, manage routings, inventory, purchase orders, and production orders. But that does not automatically mean it can create a production schedule that your factory can actually execute. In this episode, we break down the gap between ERP planning and finite production scheduling — and explain why a schedule can look perfectly reasonable in the ERP while multiple orders are competing for the same machine at the same time.
THE INFINITE-CAPACITY PROBLEM
Traditional ERP planning can place demand against resources without reserving finite blocks of actual machine time. This “infinite capacity” assumption is useful for demand and material planning, but it becomes a problem when planned dates are treated as executable shop-floor commitments. A capacity report may show that a machining centre has 40 hours of demand against only 16 available hours. It identifies the overload — but it does not decide which orders should run first, which should move, or how those decisions affect downstream operations.
CAPACITY IS MORE THAN MACHINE HOURS
Real production capacity depends on much more than a work-centre calendar. Machines have downtime. Operators have qualifications and shift patterns. Fixtures and tooling may already be occupied. Quality inspections consume resources. Maintenance removes capacity. And an eight-hour shift rarely provides eight hours of usable production time. A feasible schedule therefore has to consider the combination of machines, people, tooling, fixtures, calendars, maintenance, and process rules.
WHY SEQUENCE MATTERS
Production sequence can dramatically change the result. Running similar product families together might require only one major setup. Alternating between families can create repeated tool changes, cleaning, inspections, or fixture changes. The same orders on the same machine can therefore consume very different amounts of capacity depending on their sequence.
THE BOTTLENECK SETS THE PACE
When many orders depend on one constrained resource, keeping every upstream machine busy can actually make performance worse. More work enters the system, queues grow, WIP increases, and priorities become harder to see. Effective scheduling instead protects bottleneck capacity and controls when work is released into production.
MATERIAL AVAILABLE DOESN’T MEAN READY TO RUN
MRP may show that material exists, but that material could be under quality hold, reserved for another order, waiting for inspection, or incompatible with a specific batch requirement. Finite scheduling needs to combine material readiness with resource availability. A component arriving Wednesday only helps if the required machine also has a legal production slot when the material becomes usable.
ROUTINGS DON’T RESERVE CAPACITY
A routing tells you what comes before what. It can define cutting → machining → inspection → assembly. But a routing does not necessarily reserve the actual resource time required to execute those operations. Several orders can follow perfectly valid routings and still collide at the same machine or work centre.
WHY EXCEL KEEPS SURVIVING
This gap explains why planners continue using spreadsheets, whiteboards, notes, and local priority lists. They are combining information from ERP, MES, maintenance, quality, production, and their own shop-floor knowledge to create the schedule the factory actually follows. Excel is often not the root problem — it is the workaround for scheduling logic that exists outside the ERP.
WHAT FINITE SCHEDULING CHANGES
Finite scheduling treats production time as something that must actually be reserved. If an operation needs four hours on a machining centre, those four hours occupy a real slot. Another job cannot use the same resource during that period. The same logic can include operators, tooling, fixtures, and other required resources. When there is no legal slot, the system has to expose the conflict instead of hiding it behind another planned date.
FROM FINITE SCHEDULING TO OPTIMISATION
Once several feasible schedules exist, constraint-based optimisation can compare them. Should the plant minimise late orders? Reduce setup time? Protect bottleneck throughput? Avoid overtime? Reduce WIP? Keep the near-term schedule stable? There is rarely one universally “optimal” production schedule. The best schedule depends on the constraints the factory cannot violate and the business objectives it chooses to prioritise.
ERP VS. MES VS. APS
ERP remains essential for demand, orders, inventory, purchasing, bills of material, and transactional planning. MES provides execution truth from the shop floor. APS adds the decision layer: combining demand, materials, routings, resource availability, constraints, and current production status to create a finite, constraint-aware schedule and test alternative scenarios.
IN THIS EPISODE
You’ll learn why ERP schedules become overloaded, what infinite capacity really means, why bottlenecks and sequence-dependent setups matter, how material availability differs from production readiness, why planners fall back to Excel, and how finite scheduling and constraint-based optimisation turn production dates into an executable plan. The key idea is simple: ERP tells you what needs to happen. A production schedule has to prove when and where it can actually happen.
Become a supporter of this podcast: https://www.spreaker.com/podcast/m365-fm-a-microsoft-mvp-podcast-by-mirko-peters--6704921/support.
Production Planning Software: When Does a Manufacturer Actually Need It?
2026/09/16
When does a manufacturer actually need dedicated production planning software? It is not when the company reaches a certain size. It is not when Excel suddenly becomes “unprofessional.” And it is definitely not because a software vendor showed you a beautiful Gantt chart. The real threshold comes when production complexity, constraints, and constant change become too difficult for planners to reliably coordinate through ERP dates, spreadsheets, whiteboards, calls, emails, and experience alone. In this episode, we explore the warning signs that indicate your manufacturing planning process may have reached that point.
WHEN MANUAL PRODUCTION PLANNING STILL WORKS
Not every manufacturer needs advanced planning software. A stable factory with predictable demand, repeatable routings, relatively few shared resources, and experienced planners may work extremely well with ERP, Excel, planning boards, and direct communication. Simple tools become a problem only when the planning environment changes faster than people can reliably evaluate the consequences.
SIX WARNING SIGNS TO WATCH
We examine six signals that production planning may have outgrown its current tools: Plans change faster than people can replan. Machine breakdowns, rush orders, shortages, staffing changes, and changing customer dates create continuous replanning. Capacity exists on paper but not in reality. A machine may technically have available hours, but setups, maintenance, tooling, operator qualifications, material availability, and other constraints make that capacity unusable. The same order has different dates in different systems. ERP, MES, spreadsheets, sales, purchasing, and production may each have their own version of the expected completion date. Bottlenecks move but the plan doesn't. Today's constraint may be machining, tomorrow inspection, and next week a specific operator, tool, or downstream process. Expediting becomes the normal workflow. When almost every order becomes urgent, priorities begin replacing the production schedule. Critical dependencies live in people's heads. Experienced planners know which machines, tools, operators, routes, setups, and exceptions actually work — but the system doesn't.
ERP VS. MES VS. PRODUCTION PLANNING SOFTWARE
ERP remains essential for customer orders, bills of material, inventory, purchasing, production orders, routings, and MRP. MES provides the execution reality: what started, what finished, quantities produced, machine status, quality information, and what's happening on the shop floor. But neither automatically answers the detailed scheduling question: Given the factory as it exists right now, what work should run next — and what happens to everything else if we change the sequence? That's where dedicated production planning and scheduling software can add value.
WHAT PRODUCTION PLANNING SOFTWARE SHOULD ACTUALLY DO
A planning system should do more than display orders on a calendar. It should help planners create feasible schedules based on finite capacity, resource calendars, routing dependencies, setup requirements, material readiness, alternate resources, skills, tooling, maintenance windows, and current production conditions. More importantly, it should allow planners to test alternatives. If an urgent order moves forward, what gets delayed? If a machine goes down, where can the affected work move? If additional capacity becomes available, which orders benefit? If a supplier delivery slips, which customer commitments are now exposed? The goal isn't to remove human decision-making. It's to give planners better information about the consequences before they make the decision.
PLANNING, SCHEDULING, OPTIMIZATION AND SIMULATION
These terms are often treated as interchangeable, but they solve different problems. Production planning asks whether demand fits available capacity over a planning horizon. Scheduling determines the actual sequence of operations and resources. Optimization compares feasible alternatives against defined objectives. Simulation lets manufacturers test possible future scenarios before changing the live production plan. Understanding which problem you're trying to solve should come before selecting software.
COMPLEXITY MATTERS MORE THAN COMPANY SIZE
A large repetitive factory may have a relatively simple planning problem. A much smaller high-mix manufacturer can face enormous scheduling complexity because orders compete for shared machines, specialist operators, tools, inspection resources, alternate routings, and limited material. The need for production planning software therefore isn't primarily driven by employee count. It's driven by the number of dependencies and trade-offs planners need to evaluate — and how quickly those decisions change.
START WITH ONE DECISION
Before buying an APS platform or another planning application, identify one production decision that repeatedly causes problems. For example: When two urgent orders need the same constrained machine, can we determine which sequence is feasible and understand the delivery impact before production starts? Solve that problem first. Then expand the planning model as the organization learns which data, constraints, and rules actually matter. Because the objective isn't to automate the planner. It's to stop skilled planners from spending their day manually doing work that a planning system should already be helping them calculate. Production Planning Software: When Does a Manufacturer Actually Need It? explores where that boundary lies — and how manufacturers can recognize it before buying another platform.
Become a supporter of this podcast: https://www.spreaker.com/podcast/m365-fm-a-microsoft-mvp-podcast-by-mirko-peters--6704921/support.
Can You Simulate Your Factory Before Changing It?
2026/09/16
What happens when you change a factory before you know how the rest of the production system will react? Adding a shift, buying a new machine, reducing buffers, changing staffing, or accepting a different product mix can look like obvious solutions. But factories are interconnected systems. Improving capacity at one resource can simply move the bottleneck somewhere else. In this episode, we explore factory simulation, production planning, Digital Twins, finite capacity, bottleneck management, and manufacturing optimization — and how manufacturers can test operational changes before introducing them on the real shop floor.
WHY FACTORY CHANGES ARE HARD TO PREDICT
More machine hours do not automatically mean more customer orders shipped. An additional shift may increase machining capacity while assembly, inspection, material handling, maintenance, or qualified labor remain constrained. The result can be higher utilization at one work center while queues simply grow somewhere downstream. This is why production decisions need to consider the entire manufacturing flow, rather than optimizing individual machines in isolation.
WHAT FACTORY SIMULATION ACTUALLY MEANS
Factory simulation does not have to mean an expensive 3D visualization of an entire plant. A useful simulation models how orders move through production over time. It can represent:Routings and production sequencesMachine and labor capacityShift calendars and maintenance windowsSetup and changeover timesMaterial availabilityQueues and WIPQuality holds and inspectionsLabor skills and qualificationsBatch rulesDowntime and disruptionsDispatching and priority rulesThis allows manufacturers to test an extra shift without scheduling it, evaluate a machine before buying it, change buffer levels without disrupting production, or simulate a different product mix before customer orders are affected.
ERP VS MES VS FACTORY SIMULATION
ERP provides the commercial and production plan: demand, quantities, due dates, materials, routings, and planned capacity. MES provides evidence about what actually happened during production. Simulation adds another layer: What could happen if we change something? A work center may appear to have sufficient capacity in ERP while the real shop floor is constrained by setups, tooling, operator qualifications, material availability, inspection, or sequencing. MES history can help make simulation assumptions realistic, but historical data alone cannot answer what happens after a future shift change, capacity investment, or different dispatching rule.
FROM FACTORY DATA TO A DIGITAL TWIN
Data describes events. A factory model describes behavior. A useful Digital Twin connects products, processes, resources, people, tools, materials, quality conditions, and operating rules. It can represent the current production state and provide the starting point for testing alternative scenarios. The goal is not a perfect virtual copy of every object in the factory. The goal is a model accurate enough to support a real operational decision.
TEST CAPACITY BEFORE BUYING CAPACITY
Before investing in another machine, manufacturers can simulate alternatives such as: New machine → faster cycle time → additional shift → alternate resource → subcontracting → different sequencing rules. The important question is not simply whether machine capacity increases. It is whether throughput, lead time, queue behavior, and on-time delivery actually improve. A new machine may increase upstream output while creating an even larger queue at inspection or assembly.
PEOPLE ARE PART OF FINITE CAPACITY
Ten employees on a shift do not necessarily represent ten interchangeable units of capacity. Production may depend on specific operators who can perform setups, approve first-off parts, operate specialist equipment, or complete regulated processes. That means realistic manufacturing simulation needs to consider skills, certifications, shift coverage, supervision, support functions, and qualification constraints, not just headcount.
PRODUCT MIX AND SEQUENCING MATTER
Two production plans can contain the same number of orders and still create completely different factory loads. Different product families can require different cycle times, setups, tools, inspections, skills, and rework capacity. Average capacity figures can therefore hide the constraints that actually determine delivery performance. Sequence matters too. Grouping similar products can reduce changeovers, while prioritizing due dates may improve selected customer commitments but increase setup time. Factory simulation lets planners compare those rules against the same demand instead of relying only on averages.
START SMALL
You don't need to simulate the entire factory. Start with one decision, one constrained production area, and one measurable outcome. For example: If we add a late shift at this machining cell, can we improve due-date performance without creating an unmanageable queue at the next operation? Once the model reproduces normal production behavior credibly, alternative scenarios can be tested against that baseline. expensive capacity investment. This episode is for production planners, manufacturing leaders, operations managers, plant managers, industrial engineers, data teams, and anyone working with ERP, MES, APS, Digital Twins, or smart manufacturing.
Become a supporter of this podcast: https://www.spreaker.com/podcast/m365-fm-a-microsoft-mvp-podcast-by-mirko-peters--6704921/support.
APS Software vs. ERP: What's the Difference?
2026/09/15
Manufacturers often rely on ERP systems to manage orders, materials, inventory, purchasing, production orders, and financial transactions. But when a machine goes down, an urgent customer order arrives, or a critical operator is unavailable, a different question suddenly matters: Can the production plan actually run? That is where APS — Advanced Planning and Scheduling — differs fundamentally from ERP. ERP records and governs the business transaction. APS tests whether production can execute the plan under the real constraints of the factory. In this episode, we take a detailed look at APS software vs. ERP, why traditional ERP planning can create a false sense of certainty, and when manufacturers should consider adding finite-capacity scheduling to their production planning architecture.
ERP VS. APS: TWO DIFFERENT QUESTIONS
An ERP system provides the commercial and transactional backbone of manufacturing. It manages customer demand, bills of materials, routings, inventory, purchasing, production orders, costing, traceability, and financial records. APS looks at those same production requirements from another perspective: ERP asks: What needs to be produced? APS asks: Given our machines, people, materials, tools, calendars, setup rules, and existing workload, what can we actually produce — and when? That distinction becomes critical when several orders compete for the same constrained resources.
WHY ERP DATES AREN’T ALWAYS A FEASIBLE SCHEDULE
A production order can have a perfectly valid start date, finish date, routing, material list, and due date in ERP. That does not mean an empty machine exists at the required time. Traditional ERP and MRP planning often works with standard lead times, work-center capacity, calendars, and broader planning buckets. These assumptions are extremely useful for enterprise planning, material requirements planning, purchasing, and production control. But the real factory operates with much more specific constraints. A machine may already be occupied. The required operator may not be on shift. Material may technically be in inventory but still waiting for quality inspection. A fixture may be used somewhere else. Changing from one product family to another may require a long setup or cleaning process. This is the gap between a planned date and a feasible production schedule.
WHAT APS SOFTWARE ADDS
Advanced Planning and Scheduling software brings those physical constraints directly into the scheduling calculation. Instead of simply assigning work to dates, APS can consider finite machine capacity, resource calendars, operator qualifications, tools and fixtures, setup matrices, alternate machines, material readiness, maintenance windows, routing dependencies, campaign rules, changeovers, and production priorities. If a machine only has six available hours, a finite-capacity schedule cannot simply place ten hours of work into that shift and pretend the problem has disappeared. Something has to change. The order may move to another resource. Another order may be delayed. Overtime may be required. The production sequence may change. Or the customer promise may simply be impossible under the current constraints. APS makes those trade-offs visible before production discovers them.
FINITE CAPACITY SCHEDULING AND OPTIMIZATION
A major difference between APS and traditional production planning is finite capacity scheduling. But creating a feasible schedule is only the first step. A schedule can respect every physical constraint and still be commercially undesirable. Manufacturers may want to optimize for different objectives, including: On-time delivery, higher throughput, reduced setup time, lower work in process, bottleneck utilization, schedule stability, reduced changeovers, or protection of strategic customer orders. There is rarely one universally “best” production schedule. APS can calculate alternatives and expose their consequences. The business still has to determine which objectives matter most.
WHAT HAPPENS WHEN A MACHINE GOES DOWN?
The difference becomes particularly visible during disruptions. Imagine a critical machine fails on Monday morning while Sales simultaneously asks production to expedite an important customer order. ERP can show the production order, material availability, planned dates, purchasing status, inventory position, and customer commitment. But planners still need answers to operational questions. Can the order move to another machine? Does that machine require a different setup? Is the qualified operator available? Which existing orders would move? Would protecting this order create another late order downstream? Does the alternate machine become the new bottleneck? APS allows planners to test these scenarios against the current constraints instead of rebuilding the entire production sequence manually.
ERP, APS AND MES: WHO DOES WHAT?
ERP and APS are not the only systems involved. A useful manufacturing architecture separates business transactions, planning decisions, and execution facts. ERP answers what the customer ordered, what supply is required, and which business transactions need to be controlled. APS determines what sequence is feasible under the available resources and constraints. MES captures what is actually happening on the shop floor. This creates a closed planning loop. ERP provides demand and governed business data. APS creates the constraint-aware production schedule. MES reports actual starts, completions, downtime, quantities, scrap, and other execution events. Those execution facts can then update the planning picture, while the corresponding governed transactions flow back into ERP.
ONE SOURCE OF RECORD DOESN’T MEAN ONE SYSTEM DOES EVERYTHING
Trying to make one platform own every manufacturing decision usually creates more problems than it solves. ERP should remain the source of record for commercial and supply transactions. MES should capture execution facts close to production. APS should own the constrained scheduling view and planning scenarios. This also explains why different systems may legitimately show different dates. ERP may contain the contractual due date, APS the currently feasible completion date, and MES an expected completion based on actual production progress. The important question is not whether every date is identical. It is whether everyone understands what each date means and who owns the decision to change it.
WHY GOOD DATA MATTERS MORE THAN THE APS ALGORITHM
Installing APS does not automatically fix production planning. The scheduling engine needs accurate information about how production really operates. Resource calendars must reflect actual shifts and maintenance. Routing data must identify usable resources. Setup rules need to reflect real changeovers. Labor qualifications, tools, fixtures, material readiness, alternate resources, and other restrictions must be modeled where they materially affect scheduling. A sophisticated optimizer using inaccurate constraints simply produces an inaccurate schedule faster. This is why APS implementations often expose something important: manufacturing knowledge that previously existed only inside the planner’s head or in spreadsheets. If the same manual workaround occurs every week, it is probably no longer an exception. It has become undocumented production logic. A practical APS implementation can therefore begin with one bottleneck, one product family, or one constrained planning horizon instead of trying to model the entire factory immediately.
WHY EXCEL SURVIVES MANUFACTURING PLANNING
Excel remains popular because planners can quickly change assumptions, add missing information, and understand exactly why a calculation changed. The problem is not necessarily the spreadsheet itself. The risk appears when the spreadsheet becomes the only place where the real production logic exists. If critical setup rules, priorities, capacity assumptions, or resource restrictions live exclusively inside one planner’s workbook, the organization has created an unofficial planning system. Understanding those spreadsheets can actually be an important step toward designing a better APS implementation.
WHEN ERP SCHEDULING MAY BE ENOUGH
Not every manufacturer needs dedicated APS software. ERP scheduling may be sufficient when production routes are stable, product variety is manageable, demand is relatively predictable, capacity headroom exists, setup complexity is low, alternate routing is limited, and disruptions do not constantly force planners to rebuild the schedule. In those environments, rough-cut capacity planning, basic finite scheduling, and strong planner routines may provide enough control. Adding another specialized system would then introduce integration, data preparation, training, and operational overhead without necessarily producing enough additional value.
WHEN DEDICATED APS BECOMES IMPORTANT
The case for APS becomes stronger when multiple orders constantly compete for shared bottlenecks and the production sequence itself changes available capacity. Typical signals include frequent rescheduling, complex setups and changeovers, scarce tools or fixtures, qualified-labor constraints, alternate resources with different processing times, material timing problems, frequent disruptions, multi-stage dependencies, customer-expedite requests, or planners spending large amounts of time manually rebuilding schedules. At that point, the real requirement is not simply “better planning software.” It is the ability to calculate the consequences of a decision before committing production or promising a customer date.
Become a supporter of this podcast: https://www.spreaker.com/podcast/m365-fm-a-microsoft-mvp-podcast-by-mirko-peters--6704921/support.
Podcast reviews
Read M365.FM - Modern work, security, and productivity with Microsoft 365 podcast reviews
5 out of 5
3 reviews
★★★★★
HeyAdmin 2026/02/03
Says the right things out loud!
I really like Mirko Peters’ take on management of M365 services. His message is persuasive and articulate. Excellent information.