Skip to content
EVOLUTEME

Blog / Inside

Why the workshop shows the outline before it builds the course

A program that extracts text from a forty page manual makes errors nobody notices. So we first lay your file out by topics and write nothing until you say so.

· 7 min read

Why the workshop shows the outline before it builds the course

EvoluteMe editorial team · Examples in this article are illustrative.

Take a forty page field service procedure, the kind every service company keeps in a docx file. Ask a program to extract the text and build a course from it. We did that with our own procedure to see what would come back.

Chapter three came out fluent, confident and wrong. It told the engineer to close the ticket within three days. The procedure said three days for the first visit and fourteen for closing the ticket. Nothing crashed and nothing warned anyone. The workshop does accept your file. But because of that run, it first lays the material out by topics and goes no further until you have read the result.

What happens to a table

The procedure had a table, with deadlines on the left and actions on the right. A person looks at it and sees at once which deadline goes with which action.

When the text is extracted, the table turns into a plain sequence of cells. Sometimes they go row by row, sometimes column by column, sometimes in whatever order the layout stored them. The result looks like this: three days, seven days, fourteen days, visit the client, call, close the ticket. The deadlines are no longer attached to the actions.

That jumbled text then goes into building a course, and the model works on it diligently. It produces a task that reads perfectly well and teaches the wrong thing. The author can only catch it by reading the whole finished course line by line, which is exactly the work that someone bringing a document hoped to avoid.

Hard line breaks

The second source of junk is less obvious. In documents, lines are often broken by hand: a paragraph ends mid sentence because the page layout needed it to.

After extraction, one sentence arrives as three fragments. The model sees three pieces and treats them as three separate thoughts. It turns them into three bullet points, or joins the wrong halves and produces a sentence nobody ever wrote.

Soft hyphens inside words cause the same kind of trouble, and so do non breaking spaces and curly quotation marks. Each is trivial on its own. Together they stop the text behaving like text.

Running headers and page numbers

Every page of the procedure had the company name at the top and a page number at the bottom. Extraction put both into the text forty times.

When the model sees the same sentence forty times, it reasonably concludes that the sentence matters. The company name becomes a topic of the course. A page number ends up inside a task as a quantity, and the engineer is asked about something that was never in the procedure.

But why not just clean it up

The obvious answer is a filter, and we wrote one. It works very well until it meets the first document laid out by someone else.

A running header repeats, and so does a sentence that repeats because it matters. You cannot tell them apart without understanding the meaning, and understanding the meaning is the main job. It is a circle: to clean the text you have to understand it, and to understand it you need clean text.

A bad filter also does not fail loudly. It makes quiet mistakes. The course gets built, it looks respectable, and chapter three contains a deadline that never existed. A silent error is the worst kind in learning material, because the author is not the one who finds it. The engineer finds it, in front of a client on day three.

Other tools take a file without complaint

That is a fair point. Similar tools accept a pdf without complaining and give back something useful, and they are right to.

A summary can cope with junk. If the company name repeats forty times, the summary ignores it and gives you three paragraphs of substance. Document search can cope too: you get a quotation and look at the text around it. In both cases the reader has the last word, and the reader can see the original page.

A course cannot cope with any junk. We are not summarising the procedure. We are turning it into tasks, each with an answer that can be checked and a walkthrough. A task is a statement: this deadline, this action, this correct option. It is either true or false.

The difference is in what a mistake costs. A faulty summary gets read and dismissed. A faulty task gets learned, and the person leaves sure that they know something.

What happens instead

You bring the material in, either as pasted text or as the file itself, and press the button to lay it out into topics.

You get an outline and nothing else: the topics your document was divided into, in the order they appear. No chapter has been written and no task exists yet, so nothing has been spent on writing them. This screen costs us an extra step, and it saves the course. If a table has lost the link between deadlines and actions, you see it now, in five lines, and not three chapters into a finished course.

You move a topic, delete the one that came from a running header, and press the build button. Only then does the workshop write the chapters and the tasks, each with a walkthrough. The "material only" mode keeps everything within what you brought: it invents no facts, and it repeats deadlines and figures word for word.

What a good source looks like

Here is a practical habit that saves more time than you would expect. Before you bring the document in, look at it the way somebody who knows nothing about your subject would.

Turn the tables into sentences: "On day three, visit the client" instead of two columns. You do not have to do this. The first place where your work is really needed is the outline, where you see what the program made of your document. But preparing the document leaves the program less to get wrong, and for most documents it takes a few minutes.

Leave your headings as they are. The workshop divides the material along them, so the sections of your document become the topics of the course. The course is then organised the way you think about the subject, not the way a model decided to divide it.

Take out what does not belong in a lesson: internal file links, names of approvers, version numbers. Remove "See appendix four" too, if appendix four is not included. The model is conscientious and will build a task around an appendix it has never seen.

What you lose

The pictures. A diagram that somebody drew by hand is not carried over, and for a service procedure that is a real loss.

The formatting: nested lists, emphasis, footnotes. They carry some of the meaning, and that part stays in the document. You also lose a few minutes, because reading the outline is dull work compared with one confident click.

Why the outline comes first

We could have hidden all of this: take the file, extract the text, build the course, show the result. It makes a better demonstration, and the sales pitch is obvious: upload the file and everything will be fine.

We cannot keep that promise. A bad extraction produces no visible error, and the mistake shows up a month later, in front of a client. So the outline is the one place in the workshop where the work stops and waits for a person.

That is why the extra step stayed. Reading five lines of an outline is dull today. A course with the wrong deadlines is expensive in a month, and by then your engineers pay for it, not you.

NEXT

What to read next.

All articles

Your next step begins here.

Register