# Library to cleanup Microsoft Word HTML?

**URL:** https://discuss.prosemirror.net/t/library-to-cleanup-microsoft-word-html/3161
**Category:** Uncategorized
**Created:** [September 18, 2020, 9:46am UTC](https://discuss.prosemirror.net/t/library-to-cleanup-microsoft-word-html/3161 "2020-09-18T09:46:54Z")
**Posts on this page:** 8
**Page:** 1

<div class="post-metadata">

### Author: ![orestis](https://discuss.prosemirror.net/user_avatar/discuss.prosemirror.net/orestis/32/1707_2.png) [@orestis](https://discuss.prosemirror.net/u/orestis)
#### Post date: [September 18, 2020, 9:46am UTC](https://discuss.prosemirror.net/t/library-to-cleanup-microsoft-word-html/3161/1 "2020-09-18T09:46:54Z")

</div>

I just spent the better part of a morning trying to figure out if there’s a library out there that works in the browser and can take whatever crap HTML Microsoft Word outputs and convert it into something that has some semantic semblance.

Like, the way Word outputs list is with paragraphs that contain a bunch of inline styles, the bullet itself included (wrapped in some conditional comments). Then, to detect ordered lists you have to actually check the “bullet” itself. The better way is to actually parse some inline styles of the document and see if the list has something like `mso-level-number-format:bullet;`

This madness probably extends to other areas. Various turn-key editors (e.g. CKEditor) have some functionality to handle all this for you, but they’re very editor-specific. It would be nice if there was a library that could do this using only a plain DOMParser.

I’d think that this is something that ProseMirror users would have to deal with all the time, is there any established solution which I’m not seeing?

---

<div class="post-metadata">

### Author: ![marijn](https://discuss.prosemirror.net/user_avatar/discuss.prosemirror.net/marijn/32/15_2.png) [@marijn](https://discuss.prosemirror.net/u/marijn)
#### Post date: [September 18, 2020, 10:31am UTC](https://discuss.prosemirror.net/t/library-to-cleanup-microsoft-word-html/3161/2 "2020-09-18T10:31:32Z")

</div>

One issue is that some of the rules for such a conversion would be schema-specific. But I guess even just having a generic Word-garbage-to-clean-HTML converter (could be separate from ProseMirror, though a quick web search didn’t turn anything up) and installing that as a pasted HTML transformer, would help a lot.

Has anyone ever seen such a library? There’s the [paste-from-office plugin](https://github.com/ckeditor/ckeditor5/tree/master/packages/ckeditor5-paste-from-office) for CKEditor, which might contain a lot of the relevant logic, but that’s not open source.

---

<div class="post-metadata">

### Author: ![orestis](https://discuss.prosemirror.net/user_avatar/discuss.prosemirror.net/orestis/32/1707_2.png) [@orestis](https://discuss.prosemirror.net/u/orestis)
#### Post date: [September 18, 2020, 10:47am UTC](https://discuss.prosemirror.net/t/library-to-cleanup-microsoft-word-html/3161/3 "2020-09-18T10:47:07Z")

</div>

Yeah it’s bound to be schema-specific, and actually prose mirror already does a lot of the cleanup by throwing unknown stuff out. So even without a plugin, you get less-styled text instead of garbage, which is a good starting point.

I wonder if ProseMirror parsing rules are powerful enough to handle some stuff. Here’s an example ordered list (there’s no wrapping `ul` or anything similar):

```auto
<p class=MsoListParagraphCxSpFirst style='text-indent:-18.0pt;mso-list:l2 level1 lfo2'><![if !supportLists]><span
lang=EN-US style='mso-bidi-font-family:Calibri;mso-bidi-theme-font:minor-latin;
mso-ansi-language:EN-US'><span style='mso-list:Ignore'>1.<span
style='font:7.0pt "Times New Roman"'>&nbsp;&nbsp;&nbsp;&nbsp; </span></span></span><![endif]><span
lang=EN-US style='mso-ansi-language:EN-US'>An ordered list<o:p></o:p></span></p>

<p class=MsoListParagraphCxSpMiddle style='text-indent:-18.0pt;mso-list:l2 level1 lfo2'><![if !supportLists]><span
lang=EN-US style='mso-bidi-font-family:Calibri;mso-bidi-theme-font:minor-latin;
mso-ansi-language:EN-US'><span style='mso-list:Ignore'>2.<span
style='font:7.0pt "Times New Roman"'>&nbsp;&nbsp;&nbsp;&nbsp; </span></span></span><![endif]><span
lang=EN-US style='mso-ansi-language:EN-US'>With <u>some underlined</u> items<o:p></o:p></span></p>

<p class=MsoListParagraphCxSpLast style='text-indent:-18.0pt;mso-list:l2 level1 lfo2'><![if !supportLists]><b><i><span
lang=EN-US style='mso-bidi-font-family:Calibri;mso-bidi-theme-font:minor-latin;
mso-ansi-language:EN-US'><span style='mso-list:Ignore'>3.<span
style='font:7.0pt "Times New Roman"'>&nbsp;&nbsp;&nbsp;&nbsp; </span></span></span></i></b><![endif]><b><i><u><span
lang=EN-US style='mso-ansi-language:EN-US'>Is also nice<o:p></o:p></span></u></i></b></p>

```

- the `MsoListPargraphCxSp{First,Middle,Last}` classes seems to be reliable – so they can be used to wrap the entire list under a `ul` or `ol`.
- the content between `<!-- [if !supportLists]-->` and `<!--[endif]-->` has to be ignored, but can be used to detect the list type (`ul` or `ol`, albeit in a hacky way)
- the `mso-list:l2 level1 lfo` style can be used to detect indent level (level1, level2, level3 etc)

---

<div class="post-metadata">

### Author: ![marijn](https://discuss.prosemirror.net/user_avatar/discuss.prosemirror.net/marijn/32/15_2.png) [@marijn](https://discuss.prosemirror.net/u/marijn)
#### Post date: [September 18, 2020, 11:25am UTC](https://discuss.prosemirror.net/t/library-to-cleanup-microsoft-word-html/3161/4 "2020-09-18T11:25:12Z")

</div>

> [@orestis](#):
>
> I wonder if ProseMirror parsing rules are powerful enough to handle some stuff.

No, for much of it you’d need some kind of preprocessor.

---

<div class="post-metadata">

### Author: ![gethari](https://discuss.prosemirror.net/user_avatar/discuss.prosemirror.net/gethari/32/3168_2.png) [@gethari](https://discuss.prosemirror.net/u/gethari)
#### Post date: [September 30, 2024, 11:07am UTC](https://discuss.prosemirror.net/t/library-to-cleanup-microsoft-word-html/3161/6 "2024-09-30T11:07:46Z")

</div>

@orestis by any chance did you get some workaround over this ?

---

<div class="post-metadata">

### Author: ![prosed](https://discuss.prosemirror.net/letter_avatar_proxy/v4/letter/p/aeb1de/32.png) [@prosed](https://discuss.prosemirror.net/u/prosed)
#### Post date: [October 11, 2024, 12:26am UTC](https://discuss.prosemirror.net/t/library-to-cleanup-microsoft-word-html/3161/7 "2024-10-11T00:26:07Z")

</div>

I’m interested in implementing this - not with any sort of urgency - but the biggest problem is my lack of Microsoft Word license.

If people could just share some paste-bins or gists with problem content and an associated screenshot of what that content looks like, that would dramatically lower the bar for someone to pick this up and start running with it right away, rather than first needing a license, then needing to create some document, then getting the HTML, and finally cleaning it - they could jump straight to the cleaning part.

---

<div class="post-metadata">

### Author: ![gethari](https://discuss.prosemirror.net/user_avatar/discuss.prosemirror.net/gethari/32/3168_2.png) [@gethari](https://discuss.prosemirror.net/u/gethari)
#### Post date: [October 22, 2024, 10:05am UTC](https://discuss.prosemirror.net/t/library-to-cleanup-microsoft-word-html/3161/8 "2024-10-22T10:05:43Z")

</div>

[Office for Web](https://news.microsoft.com/microsoft365forjournalists/learning-tools/office-on-the-web-free-versions-of-your-favorite-apps/) is not FREE to use

---

<div class="post-metadata">

### Author: ![smrifat1411](https://discuss.prosemirror.net/letter_avatar_proxy/v4/letter/s/54ee81/32.png) [@smrifat1411](https://discuss.prosemirror.net/u/smrifat1411)
#### Post date: [August 18, 2026, 8:43am UTC](https://discuss.prosemirror.net/t/library-to-cleanup-microsoft-word-html/3161/9 "2026-08-18T08:43:54Z")

</div>

This exists now, if it’s still useful: **wordpaste** — [GitHub - smrifat1411/wordpaste: Clean Microsoft Word clipboard HTML on paste and keep the equations editable. OMML and MathML to LaTeX. 3.6 kB, zero dependencies, works with Tiptap, ProseMirror, Lexical or plain contenteditable. · GitHub](https://github.com/smrifat1411/wordpaste)

It is the generic converter described up-thread rather than a ProseMirror plugin: MIT, no dependencies, plain `DOMParser`, one function, string in and string out. Nothing in it knows what editor you use.

```js
import { transformPastedHTML } from 'wordpaste';

new Plugin({
  props: { transformPastedHTML: (html) => transformPastedHTML(html) },
});

```

The three list observations in this thread are exactly what it implements, so to confirm they hold up: the `MsoListParagraphCxSp{First,Middle,Last}` classes are reliable enough to bound a run, `mso-list:lN levelM` gives list identity and nesting depth, and the `<![if !supportLists]>` block has to be unwrapped rather than dropped — it is the only record of how the list was numbered. It rebuilds `<ul>`/`<ol>` from those with nesting, the original `1.` `a.` `i.` sequence, and `start` when a list doesn’t begin at 1.

One correction to the “throw unknown stuff out” approach: for lists that is worse than doing nothing. Strip the styling and the literal `1.` and `2.` stay frozen in the paragraph text, so the numbering is permanently wrong the moment anyone reorders or inserts an item. The markers have to be read before they are discarded, which means it has to happen before the schema filter, not after.

On testing without Word to hand — that was the blocker in the last reply here. The fixtures are real Word-shaped clipboard HTML with genuine `mso-` attributes and `<m:oMath>` markup, so they are reusable whatever you build: [wordpaste/test at main · smrifat1411/wordpaste · GitHub](https://github.com/smrifat1411/wordpaste/tree/main/test)

There is one thing it does that CKEditor’s paste-from-office doesn’t. Word puts an equation on the clipboard **twice** — as a rasterised screenshot and as the real OMML markup. Everyone takes the picture, so the maths is never editable again. wordpaste reads the OMML (and LibreOffice’s MathML) and converts it to LaTeX. That was the actual reason I wrote it.

Honest limits: it is **not a sanitiser** — `<script>`, inline handlers and `javascript:` URLs pass straight through. Inside ProseMirror the schema drops them, so that is fine here, but sanitise if you `innerHTML` the output yourself. Google Docs equations are already images on the clipboard, so nobody can recover those.

Playground with real Word samples, if you want to throw your own document at it before trusting it: [wordpaste playground — paste from Word into a real editor](https://smrifat1411.github.io/wordpaste/)
