What are the limits of automation and AI in content systems? In this episode, Sarah O’Keefe and Bill Swallow dive into traditional formatting, workflow challenges, and what happens when AI is introduced. The limit of automation here is not really the automation. It’s the people who say, “I don’t like the way this looks, and I’m gonna find a way to fix it. I know all sorts of ways of bending this stuff to my will.” If you’re generating content at scale, you can’t afford to do any of this stuff. Related links: Digital sovereignty in the age of AI From ad hoc to autonomous: The AI content ops maturity model Subscribe to our newsletter for updates on Sarah O’Keefe and Carlos Evia’s book: The Uninvited Author LinkedIn: Sarah O’Keefe Bill Swallow Transcript: Disclaimer: This is a machine-generated transcript with edits. Introduction with ambient background music Christine Cuellar: From Scriptorium, this is Content Operations, a show that delivers industry-leading insights for global organizations. Bill Swallow: In the end, you have a unified experience so that people aren’t relearning how to engage with your content in every context you produce it. Sarah O’Keefe: Change is perceived as being risky; you have to convince me that making the change is less risky than not making the change. Alan Pringle: And at some point, you are going to have tools, technology, and processes that no longer support your needs, so if you think about that ahead of time, you’re going to be much better off. End of introduction BS: Hi, I’m Bill Swallow. SO: And I’m Sarah O’Keefe. BS: And today we’re gonna talk a bit about automation, and more specifically what the limits are for automation. SO: Yeah, they let us out of our cages. And so here we are. And I think that in this world of I’m not going to get through the first sentence without saying AI. In this world of AI, we are suddenly facing, you know, increasing automation across every facet of everything that you could imagine. And I think it’s important to take a step back and talk a little bit about limits and what automation cannot do and maybe more importantly, why it cannot do certain things. So before we get to, you know, the AI in the room, Bill, what are the limits when we talk about generalized publishing in the pre-AI world where we’ve done a ton of formatting automation and we talk about all those kinds of things, what are some of the limits that we’ve run up against in terms of automation? BS: Yeah, automation is something that we’ve been doing quite a bit of for the past, what, twenty years, I think, or more. I’ve only been around for thirteen years. I’m thirteen years young, so… SO: Yeah, this is lies. BS: But no, a lot of our work around automation is around publishing. That is, I think first and foremost, the most common case. And with automation and publishing, you can save a lot of time and save a lot of effort, but a lot goes into building that capability because you can’t really work around edge cases or you know, fancy hand-cobbled formatting, hand page breaks, all of that thing. You have to kind of rely on the software to understand the content that you’re feeding it and it’s going to produce something that you know is a finished product: a PDF, a web page, what have you. And you know, if you deviate from the structures and conventions that it’s expecting, then all bets are off as to whether or not your automation will be successful. SO: Right. And so, from our point of view, doing lots and lots and lots of structured content work, i setting aside the tools and the technologies, what structured content really does is limit variability. Right? It limits the variability of the information or of the markup really going into your system, which then means in turn that the system can be automated and can automatically output all the outputs, all the deliverables, all the file formats that you need. And so when you have weird edge cases, and I am of course the worst offender in terms of actually finding ways to bend the software to my will, but only when I’m an author, right? As a system configuration, I’m, yeah, I think you should all follow the rules. Yeah. So here we are, and we basically say we’re gonna limit the variability of the input and standardize the input and therefore we will get standardized output. Cool. But as you said, that only works if you don’t have so what is a weird edge case or what are some of the things you’ve run into that just bollocks up automation? BS: One good case is working with content that maybe has valid structures in place. So they’re not, you know, the content itself validates against a validator. So there’s nothing wrong with it. But they’re using, let’s say, different in this case, DITA elements in a let’s say creative way. You know, so they’re you know I was on vacation last week, so that wasn’t me. SO: I see you were looking over my shoulder last week again. Yeah, that’s true. Well, what I actually ran into was I had a DITA map; it was valid, it validated, and then I used Oxygen’s validator, which you know really goes a little bit deeper and looks at things. Everything was fine, but it crashed. And eventually what I realized or what I found after some digging was that somebody who was definitely me had inserted a draft comment and the draft comment was in a table, maybe under the title, but before the table group kind of thing. It was in an unusual location and it just the processor just died. It just laid down and gave up. now I don’t know exactly whether that was because of the, you know, the core processing or something that we did in the the plugin that I was running. It doesn’t really matter. The point is it was it was valid. But it didn’t pass the processor because the processor was like, Why would anybody put a draft comment in this location? This is dumb. And then I, you know, felt berated by the processor and I moved it and then it all worked. BS: Interesting. Yeah. I would have enjoyed being there for that. SO: Yeah. Mm hmm. So anyway, we fixed it, and it was fine, and you know, and off we go. Let’s talk about pagination! BS: Pagination’s a fun one. SO: Pagination’s my favorite. I get very upset with bad line breaks or bad page breaks. BS: Yeah. And you know, we can build in rules that say, you know, only, you know, to control widows and orphans, that type of thing. You know, tell it, you know, break a table leaving X many cells. If you don’t have X many cells to break, then move the entire table to the next page, all that fun stuff. But if you have, you know, specific places where you need to break to a new page for whatever reason, that always can’t be necessarily automated, depending on what the rules are for producing your output. You know, the processor’s not going to know that, you know, you may arbitrarily want to break a page at this particular location. So you know, in many cases we cobble together a little tag that says, you know, essentially break the page here, and the processor knows to you know, when it sees that to say, okay, stop processing this page, move to the next one. And you know, it works most of the time. SO: I would never. Yeah, the problem with inserting page breaks, as I’ve learned to my great sorrow, is that, of course, later you add more content and now you have a page break a third of the way down the page because you hard-coded it in. So this is bad, and you shouldn’t do it. But if you insist on doing it, do it l as late in the process as possible. Don’t try to fix your pages when you’re 80% of the way there, because you will have to reinsert them over and over and over again. also. BS: Right. SO: Inserting empty tags that have a non-breaking space in them to introduce vertical space is wrong, and you shouldn’t do it. BS: Ha ha. I will agree with that one. Also, I mean, doing these types of things to force a page break or to force extra space, it really flies in the face of automation because technically you have to create the output in order to know where you need to insert your page breaks and then go back and add them and then automate your output again. SO: Ha ha. It looks bad. BS: Yeah. SO: So yeah. So really the limit of automation here is not really the automation, right? It’s, well, it’s me, right? It’s the people who are like, I don’t like the way this looks, and I’m gonna find a way to fix it. And I have lots more demented tricks up my sleeve. I know all sorts of ways of bending this stuff to my will. And if you’re generating content at scale, you basically can’t afford to do any of this stuff, right? It’s one thing if you’re producing, you know, one document and it’s short and/or it’s marketing content, and we’re really concerned about the appearance because of people making buying decisions. But if you’re producing, you know, ten, twenty, fifty thousand pages a year, then you just need an engine that produces this stuff. So BS: Mm-hmm. Yeah. And the page break problem is actually a good one to speak to at scale because if you’re in an environment where you’re sharing a bunch of different content and you’re reusing pieces, if you insert a page break for your own personal preference, suddenly that’s going to be in everyone else’s document that also uses that particular topic. SO: All right. So I’m hearing that page breaks are bad and I shouldn’t do them. BS: No, they’re great. It’s just that you have to be smart about it. SO: Okay. Sneak them in. Don’t get caught. Got it. So, while we’re on the subject of recalcitrant authors. Such is definitely not me. What about automation in a scenario where your content production system is not being used by the authors? BS: Yes. That’s a completely different problem. Yeah, th