M365.FM a Microsoft MVP Podcast by Mirko Peters

Mirko Peters - Founder of m365.fm, m365.show and m365con.net

M365.FM is a podcast about Microsoft 365, Microsoft Copilot, AI, Modern Work, security, governance, Power Platform, Azure, and the technologies shaping the future of work.Hosted by Microsoft MVP Mirko Peters, M365.FM brings together Microsoft MVPs, Microsoft employees, product experts, architects, developers, and community leaders from around the world.Each episode goes beyond announcements and hype to explore what Microsoft technologies mean in practice. From Microsoft 365 Copilot and AI agents to Teams, SharePoint, Power Platform, Microsoft Fabric, Entra, Purview, security, governance, adoption, and automation, M365.FM focuses on real-world experience, implementation, strategy, and lessons learned.Expect expert interviews, technical deep dives, practical explainers, and conversations with people building, implementing, and shaping the Microsoft ecosystem.If you work with Microsoft 365, Copilot, AI, Modern Work, or the Microsoft Cloud, M365.FM helps you understand what matters, what works, and what is coming next.Hosted by Mirko Peters, Microsoft MVP. Become a supporter of this podcast: https://www.spreaker.com/podcast/m365-fm-a-microsoft-mvp-podcast-by-mirko-peters--6704921/support.

  1. 1h ago

    Why Continuous Improvement Needs Better Data

    Continuous improvement is supposed to create learning that compounds over time. In many factories, however, improvement work still depends heavily on workshops, spreadsheets, isolated reports, and what people remember from the previous shift. A Kaizen event can create visible progress, but a few weeks later the same loss often appears again under slightly different production conditions. The issue is not always the quality of the improvement idea. The bigger problem is that teams often cannot prove whether the countermeasure actually worked, where it worked, and under which conditions it stopped working. WHY KAIZEN IMPROVEMENTS OFTEN DISAPPEAR A workshop creates focus for a few days. Teams map the process, identify waste, assign actions, move tools, change checklists, or adjust handoffs. Then normal production pressure returns. The supervisor has another urgent order, maintenance has another fault, planning changes the sequence, and quality puts another batch on hold. The improvement action remains somewhere in an Excel file or project tracker instead of becoming part of the operating rhythm. The important question is therefore not whether an action was completed, but whether the production condition actually improved. PDCA NEEDS A REAL CHECK STEP Plan, Do, Check, Act sounds simple, but many organizations effectively run Plan, Do, and Move On. PLAN should define a testable problem, the current condition, the expected improvement, and the hypothesis behind the countermeasure. DO means testing the countermeasure under real production conditions while recording enough context to understand what actually happened. CHECK means comparing the expected result with real production evidence. ACT means standardizing the change when the evidence supports it, or adapting, narrowing, or reversing it when it does not.    • Did the loss actually decrease? • Did the problem simply move somewhere else? • Did the change improve availability while damaging quality? • Did it work across different products, crews, and shifts? • Did maintenance, material, scheduling, or another process change influence the result? A single successful production run is not proof. Continuous improvement needs enough evidence to separate a repeatable improvement from a lucky shift. GEMBA AND DATA NEED EACH OTHER Data does not replace Gemba. Operators, supervisors, technicians, planners, and quality teams understand production conditions that systems often cannot capture. A machine record may show a ten-minute stop, while an operator knows that the stop happened because a component felt wrong, a normal material route was blocked, or the previous shift left the station in an unusual condition. At the same time, observation alone shows only one shift, one event, or one version of the problem. Connected operational data makes it possible to test whether an observation repeats across orders, products, machines, shifts, material batches, and longer periods of time.     Gemba helps teams identify where to look and which questions matter. Operational data helps test those questions across the real production pattern. Standard work then carries the learning forward so the next shift does not start from zero. START WITH THE IMPROVEMENT QUESTION One of the biggest mistakes in manufacturing analytics is beginning with the data that happens to be available. A modern production line can generate huge volumes of information from PLCs, sensors, MES systems, ERP, quality systems, maintenance applications, and spreadsheets. More data does not automatically create better decisions. A useful improvement process starts with the production question the team actually needs to answer. Instead of asking “What data do we have?”, ask “What production question are we trying to answer?” A statement such as “changeovers take too long” is still too broad. A better question would be: Why does changeover time vary significantly for the same product family on the same production line? That immediately helps define the context that matters.    • Previous product • Next product • Production order • Resource or line • Tool configuration • Material • Shift or crew • Setup start • Restart time • First acceptable unit • Stable production • Quality results • Machine alarms The goal is not to collect everything. The goal is to collect the smallest set of facts capable of changing the improvement decision. ERP EXPLAINS THE PLAN ERP provides the commercial and planning context around production. It can show planned quantities, routing, customer commitments, material requirements, order priority, and the intended production sequence. That context matters because production conditions constantly change. A countermeasure might genuinely reduce setup time while delivery performance still deteriorates because the production mix changed or planners introduced more frequent product transitions. ERP therefore helps explain what was supposed to run, in what sequence, for which demand, with which routing and materials. But ERP primarily describes intent. It does not necessarily describe what physically happened on the shop floor.   MES EXPLAINS WHAT ACTUALLY RANThe Manufacturing Execution System fills part of that gap. MES can provide the execution history behind an order, including actual operation start and finish, produced quantity, rejects, holds, rework, resources, operator transactions, material consumption, traceability, and reason codes. This allows improvement teams to connect losses with real orders and operations instead of relying on memory or daily averages.    MES data still needs interpretation. Transactions may be entered late, different crews may use reason codes differently, and a timestamp may represent when an operator confirmed something instead of the exact physical moment when it happened. The system provides evidence, but the process gives that evidence meaning. MACHINE AND IOT DATA EXPLAIN THE PHYSICAL PROCESS When teams need to understand what happened inside an operation, machine data becomes important. PLC and IoT signals can expose machine states, cycle times, alarms, speed, temperature, pressure, interlocks, motor load, restart behavior, and time to stable production. MES may show that an operation restarted at a particular time, while machine data explains what happened during the minutes before stable output returned.  • Repeated alarms • Reduced speed after restart • Temperature recovery • Manual adjustments • Interlocks • Unstable cycle times Machine data without production context can still be misleading. An alarm becomes much more useful when it can be connected to the work order, product, operation, resource, tool condition, material, and quality result surrounding the event. THE SHOP FLOOR ADDS MEANING Operators, maintenance technicians, supervisors, and quality teams fill the gaps that automated systems cannot. A reason code such as “material issue” may describe very different situations in practice. The material may have arrived late, behaved differently, carried an incorrect label, been damaged, or created a downstream problem that only became visible later. Structured codes help categorize events. Human context helps explain them. Data capture therefore needs to fit the reality of production instead of interrupting it. The useful question is not how much information an operator can enter, but what is the smallest human input that would make the next improvement decision better.     YOU NEED A SHARED FACTORY LANGUAGE ERP, MES, maintenance systems, historians, and spreadsheets can all describe the same production event differently. One system may call a resource “Line 3,” another may use “LINE03,” maintenance may track several separate assets inside the line, and a historian may still use an old PLC tag. All of these identifiers can technically be correct while still making cross-system analysis unreliable. The same problem applies to definitions. “Production complete” might mean the machine finished, MES posted the quantity, quality released the batch, the product was packed, or the order was ready to ship. A dashboard cannot decide which definition is correct. The organization has to define that shared meaning. Why Continuous Improvement Need… THE DATA MODEL BEHIND CONTINUOUS IMPROVEMENT A useful manufacturing data model preserves meaning as information moves between systems. Products, batches, work orders, resources, process steps, shifts, tools, materials, quality outcomes, maintenance conditions, production events, and timestamps need managed relationships so teams do not rebuild the same logic every time they investigate a problem. • Product • Product family • Work order • Batch • Process step • Resource • Machine or asset • Shift • Tool • Material • Quality result • Maintenance condition • Production event • Time Definitions also need ownership and versioning. Downtime, scrap, rework, good count, target rate, and changeover duration should not quietly mean different things in different reports. Products change, recipes change, tools change, machines are upgraded, routes change, and approved target rates change. Without versioning, today's process rules can accidentally be applied to historical production data. Become a supporter of this podcast: https://www.spreaker.com/podcast/m365-fm-a-microsoft-mvp-podcast-by-mirko-peters--6704921/support.

    Why Continuous Improvement Needs Better Data
  2. 2d ago

    IoT Hub Message Routing vs Event Grid — Why Telemetry and Events Are Not the Same Problem

    A temperature reading, a vibration sample, a machine cycle, and a device disconnect may all be described as events, but they do not represent the same architectural problem. In industrial IoT and manufacturing environments, confusing continuous telemetry with discrete events can create noisy workflows, incomplete production histories, unnecessary processing, and systems that react without enough context. In this episode, we break down the architectural difference between Azure IoT Hub Message Routing and Azure Event Grid by following a realistic manufacturing scenario. A press line continuously sends temperature, vibration, cycle count, energy consumption, and machine-state information through an industrial gateway. Those measurements create an operational history that engineers, data teams, maintenance teams, and production systems may need to analyze later. A device disconnect is different because it represents a change that may require another system or person to react. TELEMETRY IS A RECORD, NOT AN ALERT Telemetry represents repeated measurements over time. A single temperature value or vibration measurement usually tells you very little on its own. The real information exists in the sequence: how quickly values changed, what the machine was doing at that moment, what happened before a stop, whether measurements disappeared during a network interruption, and whether the same behavior appeared in earlier production runs. Typical industrial telemetry includes: • Temperature, vibration, pressure, energy consumption, and current draw • Machine states such as running, idle, stopped, or faulted • Cycle counts, production counters, and process measurements • Source timestamps, device identifiers, sequence numbers, and correlation information Those records may later support condition monitoring, quality investigations, energy analysis, OEE calculations, Microsoft Fabric analytics, Power BI reporting, and production optimization. That is why telemetry needs retention, replay, duplicate handling, independent consumers, and a reliable way to reconstruct the production timeline. IOT HUB MESSAGE ROUTING AS THE TELEMETRY DATA PLANE Azure IoT Hub provides the controlled device-to-cloud boundary. Devices and gateways authenticate with their own identities, send device-to-cloud messages, maintain device-management state, and can participate in controlled cloud-to-device communication. Once a telemetry message reaches IoT Hub, Message Routing determines where that data should go. Routing can inspect message properties, system properties, parts of the message body, and device twin information. This makes it possible to separate production telemetry, energy measurements, diagnostics, or other message classes before they reach downstream consumers. A common architecture might look like this: Industrial gateway → Azure IoT Hub → Message Routing → Event Hubs or Storage → Processing → Microsoft Fabric. One consumer may perform near-real-time analysis while another keeps a raw archive. A third consumer may prepare curated operational data for Microsoft Fabric. Each consumer can work independently without turning every telemetry reading into a workflow invocation. WHY ORDERING MATTERS Industrial telemetry is particularly sensitive to sequence. Imagine a machine reporting that it entered a running state, then transmitting several cycle counts, followed by a process deviation and finally a stopped state. If those records are reconstructed incorrectly, a downstream system could conclude that the machine produced parts while stopped or that a process deviation happened after production had already ended. The same problem affects downtime calculations, OEE, production counts, and condition monitoring. You therefore need to distinguish between when the source observed something, when IoT Hub received the message, and when a downstream system processed it. A robust telemetry architecture should therefore consider: • Source timestamps and cloud receipt timestamps • Stable partitioning appropriate to the asset or workload • Message IDs or sequence numbers for duplicate detection • Idempotent consumers capable of handling at-least-once delivery Network interruptions make this especially important. A gateway may buffer telemetry and send it when connectivity returns, which means Azure arrival time may be much later than the actual machine timestamp. EVENT GRID SOLVES A DIFFERENT PROBLEM Azure Event Grid is designed around publish-and-subscribe notifications. Instead of continuously reconstructing the state of a machine from thousands of readings, Event Grid tells interested systems that something changed and gives subscribers an opportunity to react. Examples in an IoT environment include device creation, device deletion, connection, disconnection, or carefully selected telemetry-derived conditions. A device-created event might trigger an asset onboarding process. A device-disconnected event might start a technical investigation. A detected engineering condition might trigger a maintenance workflow. Typical subscribers include: • Azure Functions • Logic Apps • Webhooks • Security workflows • Asset-management systems • Operational applications The fundamental architectural difference is simple: telemetry asks what an asset has been doing, while an event asks what changed and whether something should react. EVENTS SHOULD START INVESTIGATIONS, NOT DEFINE REALITY Event Grid should generally be treated as a notification mechanism, not as the final authoritative state of the physical world. Event delivery can be repeated, and events are not something a subscriber should blindly interpret as the final current state without checking. A device-disconnected event should therefore usually trigger a state check. An Azure Function might verify the current device state, inspect recent telemetry, check relevant registry information, and then decide whether the investigation should remain open. Handlers should be designed around: • Idempotent processing • Current-state verification • Event identity and timestamps • Clearly defined ownership of the resulting action This matters because duplicate events should not create duplicate tickets, duplicate alerts, or conflicting operational records. A DEVICE DISCONNECT DOES NOT MEAN THE MACHINE STOPPED One of the most important distinctions in industrial IoT is the difference between cloud connectivity, data-collection health, and production state. If an industrial gateway disconnects from IoT Hub, the cloud has lost visibility into that gateway. That does not automatically mean that the physical machine stopped. The PLC may continue controlling the equipment locally while the gateway temporarily loses connectivity. Production may continue normally, and the gateway may even buffer telemetry and upload it later. The opposite can also happen: a gateway may remain connected to Azure while its connection to the PLC or local equipment has failed. A useful architecture therefore separates cloud connection state, gateway health, machine production state, and MES work-order state. Only by combining those sources can the business determine whether a technical connectivity issue actually affected production. FILTERING IN IOT HUB VS FILTERING IN EVENT GRID Both technologies support filtering, but the purpose is different. IoT Hub Message Routing filtering answers which operational messages should enter a specific data path. Event Grid subscription filtering answers which notifications a particular subscriber should receive. For example, production telemetry may go to one Event Hubs stream while energy data goes to another destination. At the same time, a support workflow may subscribe only to disconnected events from a particular group of gateways. Confusing those two types of filtering often results in an architecture where every subscriber creates its own interpretation of the telemetry stream. A cleaner design keeps broad telemetry classification governed centrally and keeps Event Grid subscriptions focused on clearly defined reactions. THE MANUFACTURING CONTEXT LIVES OUTSIDE THE MESSAGE BUS Neither IoT Hub nor Event Grid knows what a machine means to the business. A device ID may identify a gateway, but it does not inherently know which production line the gateway belongs to, which machine it observes, which work order is running, whether a stop is planned, or whether production is already at risk. That context usually comes from multiple systems. MES provides execution context such as work orders, operations, recipes, production state, and shift information. ERP provides demand, commitments, inventory, and planning information. An asset model connects devices, gateways, PLCs, machines, cells, lines, and sites. In more advanced architectures, a digital twin or knowledge graph can help maintain those relationships. The result is a more reliable operational picture in which telemetry explains what the equipment reported, MES explains what production was doing, ERP explains why it matters, and the asset model explains how the technical components relate to the physical plant. Become a supporter of this podcast: https://www.spreaker.com/podcast/m365-fm-a-microsoft-mvp-podcast-by-mirko-peters--6704921/support.

    IoT Hub Message Routing vs Event Grid — Why Telemetry and Events Are Not the Same Problem
  3. 3d ago

    Building Azure That Survives the Real World — with Mike Martin [MVP]

    Azure architecture looks easy when everything works.The real test starts when dependencies fail, regions become unavailable, traffic spikes, assumptions turn out to be wrong, requirements change, and someone eventually asks the uncomfortable question: why did we design it this way in the first place?In this episode of M365.FM, Mirko Peters talks with Mike Martin [MVP] about what Azure architecture looks like when it has to survive real production conditions rather than just look good on a diagram.Mike brings decades of experience across development, infrastructure, architecture, leadership, coaching, training, and Microsoft Azure. One of his strongest observations is that many of the problems architects face today are not actually new. DNS still breaks. IP dependencies still matter. Costs still become a problem. Integrations still fail. Dependencies still disappear at the worst possible moment.What has changed is the level of complexity we build around those problems. FROM VISUAL BASIC TO MODERN AZURE ARCHITECTURE Mike looks back at nearly three decades in IT, starting as a Visual Basic developer in the 1990s and moving through distributed systems, networking, enterprise software, infrastructure, and eventually Azure.His key observation is simple: the industry keeps solving many of the same fundamental problems, but the architectures around them have become much more complex.Modern systems have moved from client-server applications to distributed architectures, cloud platforms, containers, microservices, Kubernetes, hybrid environments, and now AI-assisted development.That creates enormous possibilities, but it also creates a new risk: overengineering.Mike argues that many teams today make simple problems unnecessarily complicated. Good architecture is often not about adding more technology. It is about knowing what not to add. WHAT DOES AN AZURE ARCHITECT ACTUALLY DO? For Mike, architecture is not about choosing the largest number of Azure services or producing an impressive diagram.It is about understanding which components belong together, which ones should be avoided, which ones are necessary, and how to build something that remains maintainable, scalable, secure, and resilient.Architecture includes much more than compute. IdentityNetworkingData flowsSecurityMonitoringIntegrationDependenciesScalabilityOperationsRecoveryDeploymentCostMike also challenges the idea that cloud-native automatically means Kubernetes or containers.Azure provides many managed and native services that can solve problems without introducing unnecessary operational overhead.The architect’s role is to understand the complete solution and choose the simplest architecture that still satisfies the real requirements. START WITH BUSINESS REQUIREMENTS, NOT AZURE SERVICES One of the most important lessons in this episode is simple: do not start with technology.Before deciding between Azure Kubernetes Service, App Service, Azure Functions, containers, Service Bus, or another platform, teams should first understand what the solution actually needs to do.Questions to ask: Is it internal or customer-facing?How many users will depend on it?Does it need to scale globally?How long can it be unavailable?How much data can the business afford to lose?Which compliance requirements apply?What happens if the application disappears for several hours?Who is affected?What level of operational support is required?These questions lead directly into concepts such as SLAs, SLOs, RTOs, and RPOs.They also determine whether the architecture should be single-region, multi-region, active-active, active-passive, or something much simpler. RTO AND RPO WITHOUT THE BUZZWORDS RTO and RPO are often discussed as technical acronyms, but their real meaning is business-oriented. RTO — Recovery Time Objective: How quickly must the system return after a failure?RPO — Recovery Point Objective: How much data loss is acceptable?A system used for non-critical monitoring may tolerate several hours of downtime or lost data.A system supporting first responders, financial operations, commerce, or critical infrastructure may require recovery in minutes.The important point is that these numbers should not be invented by the architect.They should come from the actual business impact of failure.Once those requirements are understood, they can be translated into technical design decisions. RESILIENCE IS NOT THE SAME AS HIGH AVAILABILITYMike uses a simple analogy to explain the difference between availability and resilience.A highly available system may have another component ready to take over when the primary one fails.A resilient system is designed to absorb problems, continue functioning, recover gracefully, and return to normal without collapsing completely.That distinction matters.An application can technically be available while still providing a poor experience.It may be: SlowThrottledPartially unavailableDependent on a failing backendAffected by an integration issueUnder regional pressureReal resilience therefore requires more than uptime.It requires an architecture that can cope with load, application errors, partial failures, regional issues, broken dependencies, operational incidents, and recovery after the incident has passed. AZURE DOES NOT MAKE YOUR APPLICATION RESILIENT AUTOMATICALLY One dangerous assumption is that because Azure itself is highly available, any application running on Azure automatically inherits that resilience.It does not.Microsoft provides the services and capabilities that make resilient architectures possible.Customers still need to design for resilience.This is where architectural patterns matter: Circuit breakersLoose couplingAsynchronous processingIndependent scalingMultiple availability zonesMultiple regionsFailover mechanismsMonitoringInfrastructure as CodeTested recovery proceduresMicrosoft provides the building blocks.The architecture determines whether those building blocks actually create a resilient system. WHY ASYNCHRONOUS DESIGN MATTERS One of the strongest examples in the episode comes from systems that experience predictable traffic spikes, such as government tax portals.A common design problem occurs when every action depends on a synchronous backend operation.The user clicks a button.The frontend waits for the backend.The backend waits for another service.Another dependency slows down.Eventually the entire user experience is affected.Mike explains why loosely coupled systems are often more resilient.Instead of forcing the frontend to wait for every backend operation, applications can place work into queues or event-driven systems and allow components to process tasks independently.That creates several advantages: Different components can scale independentlyOne service can fail without taking down the whole platformBackends can process work asynchronouslyTraffic spikes become easier to absorbUser-facing components become less dependent on backend timingFailures become easier to isolateTHE MYTH OF 100 PERCENT UPTIME Another major topic in the discussion is the obsession with 100 percent availability.Mike is very clear: 100 percent uptime is not a realistic architecture target.A solution is usually composed of multiple services.Each service has its own SLA.Once several dependencies are combined, the effective availability of the complete system changes.A better architecture focuses on: Fallback scenariosRecovery proceduresRedundancyFailoverMonitoringMitigation strategiesTested recovery plansThe better question is not: how do we guarantee zero downtime?The better question is: what happens when something fails, and how quickly can the business continue operating WHEN ANOTHER NINE BECOMES TOO EXPENSIVE More availability is not automatically better.Every additional level of resilience introduces cost, infrastructure, monitoring, operational complexity, and management overhead.Whether that investment makes sense depends entirely on the business impact of failure.An internal holiday request system can probably tolerate several hours of downtime.An order-processing platform handling millions in transactions cannot.Questions that matter: How much money is lost during downtime?How many users are affected?What happens to customer trust?What happens if an API becomes slow?What happens if orders cannot be processed?What does another level of redundancy actually cost?Is the additional availability worth that cost?Reliability is ultimately an economic decision as much as a technical one. Become a supporter of this podcast: https://www.spreaker.com/podcast/m365-fm-a-microsoft-mvp-podcast-by-mirko-peters--6704921/support.

    Building Azure That Survives the Real World — with Mike Martin [MVP]
  4. 6d ago

    Why Your Production Schedule Is Wrong Before It Even Starts

    Your production schedule can look perfectly reasonable when it leaves the ERP system. Order dates line up, routing times make sense, capacity appears available, and material status suggests that production is ready to go. The problem is that the schedule is still only an assumption about the future.The moment production begins, the factory starts generating new facts. Material arrives later than expected, a batch is still waiting for quality inspection, a fixture is unavailable, an operator with a required qualification is missing, or a machine loses capacity because of a short interruption. The original schedule may have been correct when it was created, but the conditions behind it can change within minutes.This episode explores why static production schedules lose accuracy so quickly, why ERP planning is not necessarily the problem, and why modern manufacturing needs a closed feedback loop connecting planning, shop-floor execution, production data, learning, and replanning. The central argument is that a schedule is a decision about what production should try to do based on the information available at that moment. It is not a guaranteed description of what the factory will actually be able to execute. Why Your Production Schedule Is… WHY PRODUCTION SCHEDULING BREAKS DOWN Every scheduled production operation contains multiple hidden assumptions. A planned start time assumes that the previous job finishes on time, the machine remains available, the required material is usable, the correct tool or fixture is ready, and a qualified operator is present.It may also assume that setup duration remains realistic, that actual cycle time stays close to the routing standard, that quality releases the material as expected, and that another more urgent order does not suddenly compete for the same resource.That means a production schedule is not simply a table containing orders, dates, quantities, and machines. It is a collection of assumptions about future operating conditions.Typical assumptions include: Machine availabilityMaterial readinessTool and fixture availabilityOperator qualificationsSetup durationCycle timeQuality releaseResource capacityProduction sequenceCustomer prioritiesThe schedule becomes unreliable when those conditions change but the decision is not updated. ERP PLANNING IS NOT THE PROBLEM ERP remains one of the most important systems in manufacturing. It connects customer orders, inventory, bills of material, purchasing, routings, work centers, due dates, and business commitments.ERP provides the commercial intent behind production. It tells the organization what should be produced, which demand needs to be covered, which materials are required, and which customer commitments matter.The limitation appears when ERP planning is expected to understand every operational condition on the shop floor at every moment.A work center may appear available while the required fixture is still installed somewhere else. A material receipt may exist in ERP while the batch is still waiting for inspection. A person may appear on the workforce calendar while lacking the specific qualification required for the next operation.The ERP system is not necessarily wrong. The factory has simply produced newer information. PRODUCTION PLAN VS DETAILED SCHEDULE VS DISPATCH LIST One reason production planning becomes confusing is that several different decisions are often described using the same word: schedule.A production plan usually works at a broader level. It determines what demand needs to be covered, which product families should be produced, and whether enough capacity and material appear to exist across a longer planning horizon.A detailed production schedule moves closer to execution. It determines which operation should run on which resource, in what sequence, and within which time window.A dispatch decision operates even closer to the shop floor. It answers the practical question: what should this machine, operator, or work center run next based on the conditions we know right now?These decisions are connected, but they are not identical. The closer production gets to execution, the more important current operational conditions become. WHY EXCEL BECOMES THE UNOFFICIAL MANUFACTURING CONTROL SYSTEM When the official production schedule no longer matches the factory, planners frequently move into Excel. The reason is simple: Excel reacts faster.A planner can change priorities, reorder jobs, add comments, highlight material issues, record tooling problems, and send a revised sequence within minutes.That flexibility is valuable when the factory needs an operational decision immediately.The spreadsheet itself is therefore not necessarily the underlying problem. It often exposes a capability that the formal production system does not currently provide.The deeper issue appears when the production decision becomes fragmented across different places: ERP contains the original production intent.Excel contains exceptions and manual schedule changes.Supervisors hold the immediate operational sequence.Operators hold practical knowledge about what can actually run.Emails, calls, whiteboards, and shift handovers carry additional context.Once this happens, no single system represents the complete production state.The episode describes Excel as the place where exceptions, calls, and practical decisions accumulate because it can respond faster than the formal process. Why Your Production Schedule Is… THE CLOSED-LOOP PRODUCTION MODEL The solution is not simply to regenerate the schedule more often. Manufacturing needs a controlled feedback loop.A practical closed-loop production model follows five stages: PLAN — Use demand, capacity, routings, material, and business priorities to create the initial production decision.EXECUTE — Release work to the shop floor and observe what actually happens.MEASURE — Capture events that materially change the assumptions behind the plan.LEARN — Compare planned assumptions with repeated production behavior.REPLAN — Use the current production state to create the next feasible dispatch decision.This changes the objective of scheduling. Instead of trying to create one perfect schedule and defend it against reality, the organization builds a process capable of reacting when reality changes. EXECUTION PRODUCES FACTS THAT PLANNING COULD NOT KNOW Once production begins, every operation creates information that can change the remaining schedule.An operation may start late. A setup may take longer than expected. A machine may stop for twenty minutes. A quality issue may block a batch. Scrap may reduce the quantity available for the next operation. Rework may create additional demand on an already constrained resource.These are not just historical records for a weekly production report.They change what the factory can do next.For example, if an operation finishes one hour late, every downstream operation may shift. If a batch is blocked by quality, another work center may suddenly have unused capacity. If rework sends material back through an earlier operation, the rework now competes with planned production for the same machine.A useful scheduling system therefore needs execution feedback while the schedule is still active. MEASURE AT THE LEVEL OF THE PRODUCTION DECISION Manufacturing environments already generate huge volumes of machine and process data. The problem is not always a lack of data.The problem is often a lack of context.A machine reporting a twenty-minute stop tells you that capacity was lost. It does not automatically tell you which customer order was affected or which production decision should change.To support production scheduling, an event should ideally connect to information such as: Production orderOperationProductBatchResourceShiftEvent timeReason codeCurrent queueDownstream dependencyThe episode emphasizes that production measurement should operate at the same level where scheduling decisions happen: order, operation, resource, and time. Why Your Production Schedule Is…Department-level totals may tell you that output was below plan, but they cannot necessarily explain which operation consumed unexpected capacity or whether the next scheduled job is still feasible. OEE DOES NOT DECIDE WHAT SHOULD RUN NEXT Overall Equipment Effectiveness remains a useful manufacturing metric because it helps organizations understand equipment availability, performance, and quality.But OEE and production scheduling solve different problems.A machine can have strong OEE while production still misses an important customer delivery because the wrong work was placed in front of the bottleneck.Likewise, poor OEE does not automatically tell the planner what sequence should run for the rest of the shift.Production scheduling needs additional information such as actual cycle time, queue time, material readiness, tool availability, workforce qualifications, quality status, setup requirements, and customer priorities.OEE can help explain resource performance. Scheduling must determine how limited resources should be used across competing work. Become a supporter of this podcast: https://www.spreaker.com/podcast/m365-fm-a-microsoft-mvp-podcast-by-mirko-peters--6704921/support.

    Why Your Production Schedule Is Wrong Before It Even Starts
  5. Sep 24

    Microsoft 365 Copilot Without the Hype: Adoption, Governance & Getting Your Tenant Ready with Paul Keijzers [MVP]

    Microsoft 365 Copilot can help people find information and get work done faster, but its answers depend on the content and permissions already in an organization’s Microsoft 365 tenant. In this episode of the M365 FM podcast, Mirko Peters speaks with Microsoft MVP Paul Keijzers, founder of KB Works, about preparing Microsoft 365 for Copilot in a practical, responsible way. They look beyond demos and licensing to the foundations that shape Copilot results: SharePoint, Microsoft Teams, OneDrive, data quality, access controls and governance. Copilot can make existing access issues more visible. Files that have been shared broadly, outdated documents, old versions and content with unclear ownership may all affect what employees can find. Paul explains why organizations should review their sharing settings and understand where sensitive or unnecessary access exists before rolling Copilot out widely. SharePoint oversharing reports can help identify potential issues across SharePoint, Teams and OneDrive, although organizations still need to check whether flagged sharing is appropriate for their business. GOVERNANCE, LEGACY DATA AND INFORMATION PROTECTION The conversation explores how to manage years of legacy SharePoint content. Old project files may need to be retained for legal or business reasons, but that does not always mean they should appear in everyday search or Copilot results. Organizations can consider archiving content or restricting access, depending on how often it needs to be used and how the tenant is structured. Paul also highlights the need to plan for Microsoft 365 backup and recovery, including how a company can retrieve its data if it changes backup providers. Microsoft Purview is part of the wider governance picture. Paul recommends reviewing sensitivity labels, checking whether they are applied consistently, and considering automatic labeling where appropriate. Data Loss Prevention policies can help protect personal and sensitive information, while Power Automate DLP policies need attention when Copilot uses connectors to work with services such as Jira. These controls should fit the organization’s actual needs, since a global company and a small business may require different policies. SHAREPOINT STRUCTURE AND COPILOT ADOPTION Good information architecture remains important even as AI becomes better at understanding natural language. Paul discusses using metadata and content types to make information easier to organize and retrieve, while keeping SharePoint libraries practical for employees. His guideline is to avoid overly deep folder structures and excessive metadata fields, since people are less likely to maintain a system that is too complicated. Copilot can help suggest or populate information, but organizations still need a clear structure and reliable content. Copilot adoption also requires ongoing support. Rather than delivering a single training session and expecting employees to figure out the rest, organizations can share regular tips, demonstrate useful prompts and agents, and help teams solve real daily frustrations. Paul recommends starting with the work people actually do, then identifying where Copilot could save time or improve an outcome. Adoption plans should be adapted to the size and working practices of each organization instead of copied wholesale from a framework designed for a different environment. AI VALUE, AGENTS AND A PRACTICAL PATH FORWARD Mirko and Paul also discuss how organizations can assess the value and cost of AI, how agents may increasingly work alongside employees, and why administrators need ways to discover and manage agents in their environments. Paul shares his Focus Week concept in Portugal, where teams spend dedicated time working on Microsoft 365, SharePoint, governance or Copilot challenges away from their usual workplace interruptions. The central message for IT leaders is that Copilot readiness starts with understanding the Microsoft 365 tenant: who can access what, how information is organized, which content should be retained or surfaced, and how employees will learn to use AI in their work. Review sharing settings, improve information governance and connect adoption to real business needs before treating Copilot as simply another license to deploy. Become a supporter of this podcast: https://www.spreaker.com/podcast/m365-fm-a-microsoft-mvp-podcast-by-mirko-peters--6704921/support.

    Microsoft 365 Copilot Without the Hype: Adoption, Governance & Getting Your Tenant Ready with Paul Keijzers [MVP]
  6. Sep 24

    When Microsoft Messaging Looks Secure but Isn’t — Exchange Server, Hybrid and Microsoft 365 Security with Thomas Stensitzki [MVP]

    Microsoft Exchange and Microsoft 365 make it possible to run powerful messaging environments, but moving email to the cloud does not automatically make it secure. In this episode of the M365 Show, host Mirko Peters [MVP] talks with Thomas Stensitzki [MVP] about the security gaps that can hide in Exchange Server, Exchange Online, and hybrid deployments. Drawing on more than 25 years of messaging experience, Thomas explains how identity protection, careful configuration, secure mail flow, and operational discipline work together to protect an organization’s email. FROM EXCHANGE SERVER TO HYBRID AND EXCHANGE ONLINE Thomas shares how he built his career around Exchange and why email remains essential to business. He explains how a hybrid setup connects on-premises Exchange Server with Exchange Online, and why that connection needs careful planning across messaging, networking, and security teams. Microsoft 365 provides a working service with default settings, but organizations still need to configure protections such as anti-spam, anti-malware, and anti-phishing to fit their needs. IDENTITY, ADMINISTRATOR ACCESS, AND BREAK-GLASS ACCOUNTS Identity security comes first, including for service accounts and other non-human identities. Thomas discusses sensitive Exchange administrator roles, privileged access management, and why administrators should avoid using highly privileged accounts for everyday work. He also explains how to protect emergency or “break-glass” accounts, including the role of FIDO security keys and the need to plan how administrators can regain access during an outage. LEGACY SMTP, PHISHING, AND COMPROMISED ACCOUNTS Older applications and devices may still depend on basic authentication or legacy SMTP. Thomas recommends avoiding those methods where possible and describes how an on-premises relay can help route messages from systems that cannot use modern authentication. The conversation also follows a potential attack path from a malicious email to stolen credentials and unauthorized access, highlighting the value of email filtering, separate administrative accounts, and monitoring sign-in activity. MAIL FLOW, SPF, DKIM, AND DMARC Understanding the full route an email takes is essential, especially in complex environments that combine gateways, Exchange Server, Exchange Online Protection, and Microsoft Defender. Thomas explains the roles of SPF, DKIM, and DMARC in authenticating messages sent from an organization’s domain. He also recommends using dedicated subdomains for third-party services such as marketing platforms and CRM systems, and using DMARC reports to identify legitimate and suspicious senders. MICROSOFT DEFENDER AND SECURITY MONITORING Thomas discusses Microsoft Defender for Office 365 Safe Links and how link protection can help assess a URL when a user clicks it. He also covers Entra sign-in logs, suspicious sign-ins, and impossible-travel alerts. Security tools and scores can guide decisions, but administrators still need to understand what they measure, what their licenses include, and which protections their organization actually needs. CONFIGURATION, GOVERNANCE, AND OPERATIONAL DISCIPLINE The discussion moves beyond individual security settings to configuration management, least-privilege access for programmatic tools, and the risks of exposing Microsoft 365 content through APIs or agents. Thomas explains how configuration exports can help teams track changes. He also discusses Microsoft Purview sensitivity labels and data loss prevention (DLP), recommending that organizations plan their rollout carefully because these controls can be difficult to change once they are in production. BUSINESS CONTINUITY, AI, AND KEEPING EXCHANGE SECURE Backups alone may not be enough if an organization loses access to its Microsoft 365 tenant, domains, or configuration. Thomas stresses the importance of planning for business continuity before an incident occurs. He also considers how AI may help both defenders and attackers, and describes his consulting work helping organizations update Exchange Server environments and move to Exchange Online or hybrid configurations. RAPID-FIRE QUESTIONS AND FINAL SECURITY ADVICE In the rapid-fire round, Thomas chooses Exchange Server, a long-term hybrid architecture, and PowerShell. He also discusses Exchange Server certificate management and recommends Manfred Huber as a future guest. His closing advice is straightforward: keep Exchange environments up to date, follow the Exchange Product Group blog, apply patches, watch for default configuration changes, and prepare for the deprecation of Exchange Web Services in Exchange Online. Become a supporter of this podcast: https://www.spreaker.com/podcast/m365-fm-a-microsoft-mvp-podcast-by-mirko-peters--6704921/support.

    When Microsoft Messaging Looks Secure but Isn’t — Exchange Server, Hybrid and Microsoft 365 Security with Thomas Stensitzki [MVP]
  7. Sep 18

    Content Understanding & Document AI - Simply Explained

    Invoices, contracts, receipts, forms, scanned PDFs, and email attachments contain valuable business information — but most automation still struggles to turn those documents into reliable, structured data. In this episode of M365 FM – Simply Explained, we break down Microsoft Content Understanding, Document AI, OCR, AI Builder, Power Automate, Dataverse, and Power Apps and explain how they work together to transform documents into usable business data and automated processes. You’ll learn why traditional OCR is only the beginning. OCR can recognize text such as an invoice number or amount, but it does not automatically understand whether a number represents an invoice ID, purchase order, bank account, tax value, or phone number. Document AI adds context, structure, and business meaning to extracted information. We explain how Microsoft Content Understanding can take documents and images, extract defined fields, classify information, and return structured results that downstream systems can use. Instead of asking AI to summarize an entire document, organizations can define a schema containing fields such as supplier name, invoice date, invoice number, invoice total, document type, and line items. The episode also shows how Microsoft Power Platform turns document extraction into a complete business workflow. Power Automate can detect new documents in email or SharePoint, send them for extraction, validate the returned information, create records, trigger approvals, and route exceptions to the right person. Dataverse can store structured document records and process history, while Power Apps can provide a human review interface for correcting uncertain or missing information. We also cover one of the most important parts of Document AI: confidence scores and human-in-the-loop review. A high confidence score does not automatically mean a value should be trusted. Critical information such as invoice totals, payment instructions, bank details, or contract dates may still require additional validation against business rules and existing systems. You’ll discover how validation can check whether suppliers exist, purchase orders match, totals make sense, dates are valid, and duplicate invoices have already been processed. This combination of AI extraction, validation rules, governance, and human review is what turns Document AI from an impressive demo into a reliable business process. We also look at practical first use cases including invoice processing, employee onboarding forms, claims, contract expiry dates, supplier documents, procurement workflows, HR documents, legal documents, service requests, and customer forms. The key is to start with one document type, one clear decision, a defined owner, and a review path for exceptions. By the end of this episode, you’ll understand the complete Document AI pattern: Document → Content Understanding → Structured Data → Validation → Power Automate → Dataverse → Human Review → Business Action The goal is not simply to process more PDFs. It is to stop people searching through documents for basic information and instead move structured, validated data directly into the business process where decisions happen. Subscribe to M365 FM for practical episodes about Microsoft Content Understanding, Power Platform, Power Automate, AI Builder, Microsoft AI, automation, Copilot, document processing, and the future of work. Become a supporter of this podcast: https://www.spreaker.com/podcast/m365-fm-a-microsoft-mvp-podcast-by-mirko-peters--6704921/support.

    Content Understanding & Document AI - Simply Explained
  8. Sep 18

    IoT Hub Message Routing vs Event Grid — Why Telemetry and Events Are Not the Same Problem

    A machine sends temperature readings, vibration data, cycle counts, power consumption, and operating states every few seconds. Then the gateway suddenly disconnects. Are all of those messages simply “events”? Technically, you could describe them that way. Architecturally, that can create serious problems. Azure IoT Hub Message Routing and Azure Event Grid solve different problems. One path is designed around preserving and distributing operational data. The other is designed around notifying systems that something changed and may require a response. Treating them as interchangeable can leave you with expensive workflows processing routine sensor data—or important signals buried inside a telemetry pipeline nobody is actively watching. In this episode of M365 FM, we follow a manufacturing machine through a real Azure IoT architecture and explain where IoT Hub, Event Grid, Event Hubs, Microsoft Fabric, Power BI, MES, ERP, Functions, and Logic Apps actually belong. WHAT YOU WILL LEARN In this episode, we explore: Why machine telemetry and discrete business or lifecycle events require different architecture patternsHow Azure IoT Hub Message Routing works as part of a telemetry data planeWhere Azure Event Grid fits into event-driven and reactive architecturesWhy message ordering matters for manufacturing telemetryWhy Event Grid should not become your primary high-volume telemetry busWhy IoT Hub routing should not be forced into every notification workflowHow IoT Hub and Event Grid can work together in the same architectureHow Event Hubs can support independent stream-processing consumersWhy raw telemetry should often be retained for traceability and later investigationHow Microsoft Fabric and Power BI can consume prepared operational dataWhy MES, ERP, and asset models provide context that device data alone cannot provideHow device disconnect events should be interpreted without automatically assuming production stoppedHow duplicate delivery, retries, timestamps, and idempotency affect reliable industrial architecturesHow to design condition monitoring, predictive maintenance, quality traceability, and production-disruption workflowsHow to decide whether a message belongs on the data plane, the response path, or bothTELEMETRY IS A RECORD OVER TIME Telemetry is not valuable because one temperature reading arrived. It becomes valuable because thousands of readings together describe what happened. A production machine may continuously report: Temperature and vibration measurementsMotor current and energy consumptionCycle counts and production countersRunning, idle, stopped, or faulted statesSource timestamps and sequence informationDiagnostic and equipment-health informationA single temperature value might mean very little. The sequence around that reading tells the story. Was the machine warming up? Was it already producing? Was vibration increasing at the same time? Did cycle time begin to increase? Did the machine stop shortly afterward? Telemetry therefore needs a path designed around sequence, retention, replay, independent consumers, and traceability. EVENTS EXIST TO START A RESPONSE An event serves another purpose. An event says: Something changed. A system or person may need to react. Examples include: A new device was registeredA gateway disconnected from IoT HubA device reconnectedA device was deletedA monitoring process detected a condition requiring investigationAn inspection completed and another workflow can beginThe recipient usually does not need hours of telemetry before starting the first step. It needs enough information to identify what happened and determine the appropriate response. That response might involve: Starting an Azure FunctionTriggering a Logic AppOpening a support investigationUpdating an asset recordChecking the current device stateCalling an external application through a webhookNotifying the team responsible for the affected systemThe event starts the investigation. It does not necessarily contain every fact needed to make the final operational decision. WHY “EVERYTHING IS AN EVENT” BREAKS DOWN Sending every sensor measurement into event-triggered workflows can look attractive during a proof of concept. Then production scale arrives. Every reading triggers another Function. Another Logic App evaluates something. Another integration receives another message. Maintenance creates its own subscription. Quality creates another. Energy management creates another. Soon, every team has slightly different filtering, state management, retry handling, and storage logic. A temperature measurement is not automatically an incident. It may contribute to an incident later, but the continuous measurements should remain available as evidence. When routine telemetry starts generating constant notifications, users can also begin ignoring alerts because the system has trained them to expect noise rather than actionable information.  WHAT AZURE IOT HUB ACTUALLY DOES Azure IoT Hub provides the device-facing cloud boundary. It supports areas such as: Secure device identitiesDevice-to-cloud messagingCloud-to-device communicationDevice managementDevice twinsControlled access between connected equipment and Azure servicesIoT Hub knows that an authenticated device sent a message. It does not automatically understand what that message means to production. A gateway might report: state = running But “running” could mean: Producing approved partsDry cyclingRunning setupPerforming reworkMoving without materialThat context usually comes from additional systems such as the MES, ERP, asset model, or production application. Become a supporter of this podcast: https://www.spreaker.com/podcast/m365-fm-a-microsoft-mvp-podcast-by-mirko-peters--6704921/support.

    IoT Hub Message Routing vs Event Grid — Why Telemetry and Events Are Not the Same Problem

Ratings & Reviews

5
out of 5
3 Ratings

About

M365.FM is a podcast about Microsoft 365, Microsoft Copilot, AI, Modern Work, security, governance, Power Platform, Azure, and the technologies shaping the future of work.Hosted by Microsoft MVP Mirko Peters, M365.FM brings together Microsoft MVPs, Microsoft employees, product experts, architects, developers, and community leaders from around the world.Each episode goes beyond announcements and hype to explore what Microsoft technologies mean in practice. From Microsoft 365 Copilot and AI agents to Teams, SharePoint, Power Platform, Microsoft Fabric, Entra, Purview, security, governance, adoption, and automation, M365.FM focuses on real-world experience, implementation, strategy, and lessons learned.Expect expert interviews, technical deep dives, practical explainers, and conversations with people building, implementing, and shaping the Microsoft ecosystem.If you work with Microsoft 365, Copilot, AI, Modern Work, or the Microsoft Cloud, M365.FM helps you understand what matters, what works, and what is coming next.Hosted by Mirko Peters, Microsoft MVP. Become a supporter of this podcast: https://www.spreaker.com/podcast/m365-fm-a-microsoft-mvp-podcast-by-mirko-peters--6704921/support.

You Might Also Like