feat(export): native DOCX export via html-to-docx (opt-in) (#7568)

* feat(export): native DOCX export via html-to-docx (opt-in)

Addresses #7538. The current DOCX export path shells out to LibreOffice,
which means every deployment that wants a Word download either installs
soffice (~500 MB) or loses that export. This PR adds a pure-JS
alternative: render the HTML via the existing exporthtml pipeline, then
feed it to the `html-to-docx` library in-process to produce a valid
.docx buffer — no soffice required, no subprocess spawn, no temp file
dance for the DOCX case.

Behavior:
- `settings.nativeDocxExport` (default `false`) gates the new path so
  existing deployments see zero behavior change.
- When enabled, `type === 'docx'` requests skip the LibreOffice branch,
  run `html-to-docx(html)`, and return the buffer with the
  `application/vnd.openxmlformats-officedocument.wordprocessingml.document`
  content-type.
- If the native converter throws, the handler falls through to the
  existing LibreOffice path — so flipping the flag on is safe even on a
  mixed-installation where soffice is still present as a backstop.
- Other export formats (pdf, odt, rtf, txt, html, etherpad) are
  unchanged.

Files:
- `src/package.json`: `html-to-docx` dep (pure JS, no binary reqs)
- `src/node/handler/ExportHandler.ts`: new DOCX branch gated on the
  setting, with fall-through on error
- `src/node/utils/Settings.ts`, `settings.json.template`,
  `settings.json.docker`, `doc/docker.md`: wire up the new setting +
  env var (`NATIVE_DOCX_EXPORT`)
- `src/tests/backend/specs/export.ts`: two new tests — asserts the
  exported buffer is a valid ZIP (PK\x03\x04 signature) and the
  response carries the correct content-type — both with
  `settings.soffice = 'false'` to prove the path doesn't need soffice
  at all.

Out of scope for this PR:
- Native PDF export (would need a PDF rendering step — separate
  undertaking, and the issue acknowledges the `pdfkit`/puppeteer size
  trade-off).

Closes #7538

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(7538): skip native DOCX test when html-to-docx isn't installed

The upgrade-from-latest-release CI job installs deps from the previous
release's package.json (before this PR adds html-to-docx) and then
git-checkouts this branch's code without re-running pnpm install.
Under that one workflow the new test can't find the module and fails
on the LibreOffice fallback, masking that the native path actually
works in every normal install.

Guard the describe block with require.resolve('html-to-docx'); Mocha's
this.skip() on before cascades to the sibling its. Regular backend
tests (pnpm install against this branch's lockfile) still exercise it.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(7538): spec for soffice-free DOCX/PDF export and DOCX import

Captures the agreed scope expansion of PR #7568: replace the flag-gated
native DOCX path with a soffice-first selection cascade, add native PDF
export via pdfkit + a small htmlparser2-driven walker, and add native
DOCX import via mammoth. Also defines a shared HTML sanitizer
(stripRemoteImages) used by both export converters to close the
SSRF surface that Qodo flagged on the html-to-docx path.

The spec drops the nativeDocxExport setting and its env var; with
soffice configured, behavior is unchanged, and with soffice null,
docx/pdf export and docx import all work in-process. odt/doc/rtf
(and pdf import) keep needing soffice and are documented as such.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(7538): implementation plan for native DOCX/PDF and DOCX import

Bite-sized TDD task breakdown of the soffice-free export/import work:
rebase, deps, sanitizer, PDF walker, mammoth wrapper, ExportHandler
cascade, route guard, ImportHandler branch, UI fix, flag rollback,
verification + Qodo reply.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* chore(7538): add pdfkit, htmlparser2, mammoth deps

Pure-JS, no native binaries:
- pdfkit ^0.18.0  (PDF rendering)
- htmlparser2 ^12 (SAX parser used by walker + sanitizer)
- mammoth ^1.12   (DOCX -> HTML for native import)
- @types/pdfkit ^0.17 (dev)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(7538): add stripRemoteImages HTML sanitizer

Drops <img src=> elements pointing at non-data, non-relative URLs to
prevent the DOCX/PDF converters from making outbound requests via
plugin-modified HTML. Closes Qodo finding #4 against the
html-to-docx path; will be wired into both export branches in
the cascade refactor.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(7538): native PDF export via pdfkit + htmlparser2 walker

Renders pad HTML to a PDF Buffer in-process: headings, paragraphs,
lists, links, inline emphasis, data:-URI images. Remote images are
explicitly skipped at the walker (defense-in-depth on top of the
shared stripRemoteImages sanitizer).

PDFs are emitted with compress:false so accessibility/SEO indexers
that don't FlateDecode can still extract text. Pads are small enough
that the size cost is negligible.

Walker is 167 LOC, well under the spec's 500-LOC bail-out
threshold for switching to pdfmake+html-to-pdfmake+jsdom.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(7538): native DOCX import via mammoth

Wraps mammoth.convertToHtml so a soffice-less Etherpad can ingest
.docx files. Images are coerced to data: URIs at the converter
boundary so the import pipeline never sees a remote src=.

Includes a tiny generated DOCX fixture (heading, paragraph, list)
under tests/backend/specs/fixtures/ for both this wrapper test and
the upcoming end-to-end import test.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(7538): soffice-first cascade in ExportHandler

Replaces the flag-gated DOCX branch with a deterministic dispatch:
soffice if configured, native DOCX/PDF otherwise, 5xx on native
error. Both native paths run plugin-modified HTML through
stripRemoteImages first.

Test changes:
- existing native DOCX block now sets soffice=null (was 'false', a
  truthy non-null string that sidestepped the route guard); fixes
  Qodo finding #3.
- new native PDF integration tests assert %PDF- header and
  application/pdf content-type with soffice=null.
- new negative test: with soffice=null, /export/odt still returns
  the 'not enabled' message.
- the legacy 500-on-export-error test now uses /bin/false so it
  exercises the soffice error path explicitly (the cascade dropped
  the ad-hoc 'false' string; .doc has no native path so this still
  works as a soffice error probe).

Integration tests for native DOCX/PDF currently fail because the
/export route guard still treats both formats as soffice-only;
the next commit fixes that.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(7538): allow docx/pdf through export guard without soffice

Tightens the no-soffice block to ['odt','doc'] only — formats with
no native path. docx and pdf are handed to ExportHandler, which
dispatches to the in-process converters. Closes Qodo finding #2.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(7538): native DOCX import path in ImportHandler

When soffice is null and the upload is .docx, run mammoth and feed
the resulting HTML through setPadHTML. Other office formats
(pdf/odt/doc/rtf) are explicitly rejected with uploadFailed instead
of silently falling through to the ASCII-only path.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(7538): always show DOCX/PDF export links

Native paths (#7538) make DOCX and PDF available regardless of
soffice presence, so unconditionally render those links. ODT still
gates on exportAvailable. Closes Qodo finding #2 on the UI side.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* refactor(7538): drop nativeDocxExport flag

Selection is now purely soffice-presence-driven (cascade in
ExportHandler). The opt-in setting and its NATIVE_DOCX_EXPORT env
var are no longer needed -- soffice configured means soffice path;
soffice null means native path (DOCX, PDF, and DOCX import).
Reverts the additive surface introduced earlier in this PR.

Also updates the SOFFICE doc row to reflect that null no longer
means 'plain text and HTML only' -- docx/pdf export and docx
import now work natively without soffice.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(7538): tighten link annotation assertion

CodeQL flagged the loose 'raw.includes("etherpad.org")' as
'incomplete URL substring sanitization' (a false positive in test
context, but worth fixing). Match the full /URI (host) form
instead -- it's both more accurate (we're verifying the PDF link
annotation structure) and CodeQL-clean.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(7538): cleaner DOCX/PDF output + round-trip test coverage

DOCX:
- New extractBody helper drops <head>/<style> and the leading
  newline inside <body> so html-to-docx doesn't render CSS or
  prefix paragraphs with empty space.
- New wrapLooseLines pre-processor wraps loose pad lines in <p>
  before the converter sees them. html-to-docx renders <br>
  outside <p> as a new <w:p> (full empty line in Word); inside
  <p> it correctly emits <w:br/> (soft break). Etherpad's HTML
  uses bare <br> for every line, so this was making single
  Enters look like double Enters in the Word output.

PDF:
- Walker SKIP_TAGS rejects head/style/script/title/meta/link
  content -- prior version dumped CSS into the rendered PDF.
- New breakLine() helper combines flushLine() with moveDown(1).
  pdfkit's text('', false) closes the continued run but does
  NOT advance the cursor, so consecutive runs were stacking at
  the same y-coordinate. <br>, end-of-block, and list items
  now use breakLine().
- ontext collapses runs of whitespace and drops pure-whitespace
  text nodes so pretty-printed source HTML doesn't render its
  formatting newlines.

Round-trip:
- New backend test: pad text -> DOCX export -> DOCX import ->
  new pad. Asserts content survives the trip.
- New PDF sanity test: extracts visible text from the PDF stream
  and asserts the source pad text appears verbatim.
- 6 new unit tests for extractBody and wrapLooseLines plus 1 for
  PDF walker SKIP_TAGS coverage.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(7538): CodeQL ReDoS + import.ts type error

- BR_PARA_RE was /(?:\s*<br>\s*){2,}/ -- two adjacent \s* runs can
  match the same chars, so on '<br>\t<br>\t<br>...' the regex
  backtracks exponentially. Re-anchored to match a fixed first <br>
  followed by one or more additional <br>s, so each whitespace run
  has exactly one home.
- import.ts: fetchBuffer was typed Promise<Buffer> but call sites
  chained .expect(200) on it, which only works on supertest's Test
  object. Return the Test (typed any) so the chain is preserved.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(7538): plugin-aware HTML cleanup, PDF text-align, monospace

ep_headings2/ep_align emit one heading-styled blank-line block after
every styled line in the pad ('<h1 style=text-align:right></h1>'),
which both html-to-docx and our pdfkit walker render as a full empty
paragraph. Plus the pdfkit walker had no support for text-align or
monospace, so right/center alignment and 'code' lines rendered the
same as plain body text.

- New dropEmptyBlocks helper strips empty h1-h6/p/code/pre/div/
  blockquote wrappers in preprocessing. Iterates so nested empties
  collapse too. Applied before both DOCX and PDF conversion.
- PDF walker now reads style='text-align:left|center|right|justify'
  on block elements (h1-6, p, div) and passes it as pdfkit's align
  option. align is applied once per continued run, then reset on
  flushLine so the next block can pick up its own value.
- PDF walker handles <code>, <tt>, <kbd>, <samp> as inline monospace
  (Courier) and <pre> as block monospace (Courier + breakLine on
  open/close).

11 new unit tests:
- 4 for dropEmptyBlocks (heading wrappers, code, nesting,
  pass-through)
- 1 for PDF text-align (compares the BT matrix x for left vs right)
- 2 for Courier in <code> and <pre>

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(7538): separate adjacent heading-style blocks on import

Round-trip bug: ep_headings2 emits <h1>/<h2>/<code> in pad HTML.
Mammoth round-trips them as adjacent <h1>A</h1><h2>B</h2><p>C</p>
on import. Etherpad's server-side content collector has a default
_blockElems set of just {div, p, pre, li}, and ep_headings2 only
registers the CLIENT-side aceRegisterBlockElements -- not the
server-side ccRegisterBlockElements. So h1/h2/code end up being
treated as inline by the importer, and adjacent blocks merge into
a single pad line.

Fix: insert <br> after </h1>...</h6>/</code> when followed by
another block. Server-side workaround keeps this PR self-contained
regardless of plugin version. The right long-term fix is to extend
ep_plugin_helpers' lineAttribute factory to register both hooks
(filed as a follow-up).

Tests:
- 5 unit tests for separateAdjacentHeadingBlocks
- New end-to-end round-trip test asserts H1+H2+P land on three
  separate pad lines after the import path.

Plus the prior PDF text-align/Courier/code commit also included
here:
- code/tt/kbd/samp inherit text-align from style attribute
- pre inherits text-align too

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(7538): blank-line round-trip + DOCX code monospace + a==c tests

DOCX round-trip dropping blank pad lines:
- wrapLooseLines now emits an explicit <p></p> marker for each blank
  line in a <br> run, instead of collapsing all gaps into a single
  paragraph break. (N consecutive <br>s -> 1 paragraph boundary +
  N-2 empty <p></p> markers, mapping to N-1 blank pad lines.)
- mammoth's docxBufferToHtml now passes ignoreEmptyParagraphs:false
  so the empty <w:p> entries survive the import side. mammoth's
  default of true was silently dropping them.
- dropEmptyBlocks no longer strips <p></p> -- that's the meaningful
  marker for the round-trip. Empty <h1>/<code>/<pre>/<div>/
  <blockquote> are still stripped (plugin noise).

DOCX <code> rendering as monospace:
- New applyMonospaceToCode wraps code/tt/kbd/samp/pre content in a
  <span style="font-family:'Courier New', monospace">. html-to-docx
  honors that and emits <w:rFonts w:ascii="Courier New".../>, which
  Word renders as Courier. The bare <code> tag is otherwise just
  a no-op for html-to-docx.
- Applied only on the DOCX export path (PDF walker already handles
  monospace via Courier font selection).

Round-trip tests:
- New a==c suite: txt, etherpad, html, docx -- export from src,
  import to dst, re-export and compare against the meaningful
  invariant (line text for binary formats; trimmed body for HTML).
- HTML test tolerates one trailing <br> per round-trip because
  setPadHTML appends a final <p> on import; this is pre-existing
  core behavior, not our bug.
- DOCX test normalizes trailing newline run (same reason).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(7538): preserve <w:jc> alignment through mammoth round-trip

mammoth doesn't expose Word's paragraph alignment (`<w:jc>`) when it
converts a docx to HTML -- there's no equivalent in its style-mapping
machinery. To keep alignment through DOCX round-trips we walk the
docx's document.xml directly, pull the `w:val` from each `<w:p>`'s
`<w:jc>`, and inject `style="text-align:..."` onto the matching
block element in mammoth's output by document order.

Word's w:jc accepts more values than CSS text-align; we map left/
start, center, right/end, both/justify/distribute and skip the rest
(start/end take left/right because we don't track ltr/rtl from the
docx for now).

Combines with the upstream ep_align PR (ether/ep_align#183) for the
full round-trip: this PR makes the mammoth output carry the
alignment style; ep_align#183 makes the importer pick it up.

Closes the alignment side of #7538.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(7538): preserve <a href> inside <code>/<pre> in DOCX export

html-to-docx silently drops <a href> children of <code>/<pre> tags
(and of styled <span>s, but the <code> wrapper is the active offender
here). The pad-export HTML produced by ep_headings2 + ep_align uses
<code style='text-align:right'>...<a href>...</a>...</code> for each
'Code'-style line, which lost its links on every DOCX export.

Workaround: applyMonospaceToCode now drops the code/pre/tt/kbd/samp
wrapper entirely. The non-anchor content gets wrapped in monospace
spans; anchors are emitted unstyled so they keep their hyperlink. For
block-level usage (<pre>, or <code> with an inline style attr) we
emit a wrapping <p> and forward the text-align style. Run BEFORE
wrapLooseLines so the <p> doesn't get double-wrapped.

Tests added:
- inline <code> -> just a styled span (no <code> wrapper)
- <code style='text-align:right'> -> <p style> wrap
- <pre> -> always block-wrapped
- <tt>/<kbd>/<samp> -> inline span only
- regression: <a href> inside <code> survives html-to-docx round-trip
  with both the URL in word/_rels/document.xml.rels AND a <w:hyperlink>
  in the document body

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(7538): HTML import line-break drift between blocks

Etherpad's HTML export wraps each pad line in <p>...</p> (or <h1>,
<code>, etc.) and then appends a <br> between lines. The closing
block tag already ends the line for contentcollector, so the trailing
<br> is redundant -- and on import the server collector counts BOTH
as line breaks, doubling every blank line between paragraphs and
inserting an extra blank between adjacent headings.

Two fixes, both gated on the runtime block-element registry so they
don't double-trigger when the underlying plugin already handles
adjacency:

1. HTML import path now runs the new collapseRedundantBrAfterBlocks
   helper before setPadHTML. Drops a single <br> immediately
   following </p>/</h1-6>/</code>/</pre>/</div>/</blockquote>/</ul>/
   </ol>/</li>/</table>/</tr>/</td>/</th>. Multiple consecutive
   <br>s after a block keep all but the first (the rest still
   represent intentional blank lines).

2. The DOCX-import separateAdjacentHeadingBlocks workaround now
   checks whether 'h1' is in the runtime ccRegisterBlockElements
   set before inserting <br>s. When ep_headings2 has the new server
   hook (per ep_plugin_helpers#14 + the upcoming ep_headings2 PR),
   the workaround correctly stays out of the way -- otherwise it
   adds an extra blank line per heading transition.

Also fixed a subtle ts-check failure on the import.ts test changes
and a leftover implicit-any in ImportDocxNative's alignment
preserver.

Tests added:
- collapseRedundantBrAfterBlocks: 5 unit tests (each block tag,
  whitespace tolerance, multiple <br> keeping intentional blanks)
- HTML import: 'does not introduce a blank line between H1 and H2',
  'preserves blank-line count between H1 and H2 (realistic shape)'
  reproduces the 5-blanks-where-2-expected bug from the user's
  round-trip pad.

1054 backend tests pass locally (the 6 failures are the pre-existing
favicon/webaccess send@1.x dotfile-path issue from running under
.claude/, doesn't reach CI).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(7538): skip plugin-dependent HTML import tests on no-plugin CI

The two new HTML-import-adjacency tests assume ep_headings2 (or
another plugin) has registered h1/h2 as server-side block elements
via ccRegisterBlockElements. Without that, contentcollector treats
<h1>/<h2> as inline and adjacent ones merge into a single pad line
-- making the assertions inapplicable.

CI's backend-tests job runs without plugins installed, so guard
the describe block with a runtime hooks.callAll() check and skip
when h1 isn't a registered block. Local dev with ep_headings2 (and
the local plugin patch wiring ccRegisterBlockElements) still
exercises both tests.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
This commit is contained in:
John McLear 2026-05-09 01:33:50 +08:00 committed by GitHub
parent afee7969dd
commit c47ffd5705
No known key found for this signature in database
GPG key ID: B5690EEEBB952194
15 changed files with 4191 additions and 35 deletions

View file

@ -90,7 +90,72 @@ exports.doExport = async (req: any, res: any, padId: string, readOnlyId: string,
return;
}
// else write the html export to a file
// Soffice-first dispatch (issue #7538). When soffice is configured
// we keep the legacy convert-via-tempfile path; when it's not, we
// hand DOCX to html-to-docx and PDF to our pdfkit walker — both
// pure-JS, in-process. No fallback chain: native errors surface as
// 5xx so admins see real failures instead of silent shadowing.
const {sofficeAvailable} = require('../utils/Settings');
const sofState = sofficeAvailable();
const goNative = sofState === 'no'
|| (sofState === 'withoutPDF' && type === 'pdf');
if (goNative) {
const {
stripRemoteImages, extractBody, wrapLooseLines, dropEmptyBlocks,
applyMonospaceToCode,
} = require('../utils/ExportSanitizeHtml');
// The HTML pipeline returns a full document (head, style, body); the
// legacy soffice path renders that fine, but the in-process
// converters need just the body content to avoid leaking CSS into
// the output and to drop the document-level whitespace that creates
// stray paragraph breaks at the top of the result.
// dropEmptyBlocks strips heading-styled blank-line wrappers that
// ep_headings2 emits between every styled line.
const bodyHtml = dropEmptyBlocks(stripRemoteImages(extractBody(html)));
html = null;
try {
if (type === 'docx') {
// applyMonospaceToCode strips `<code>`/`<pre>`/`<tt>` wrappers
// (html-to-docx ignores them AND has a bug where it drops
// `<a href>` children of those tags) and emits styled
// monospace spans, forwarding any block-level alignment style
// to a wrapping `<p>`. Run BEFORE wrapLooseLines so the
// resulting `<p>` lands at the loose-line boundary instead
// of getting double-wrapped.
//
// wrapLooseLines then handles `<br>` semantics: bare `<br>`
// outside `<p>` becomes a soft break, `<br><br>` becomes a
// paragraph boundary plus blank-line markers.
const docxHtml = wrapLooseLines(applyMonospaceToCode(bodyHtml));
const htmlToDocx = require('html-to-docx');
const buf = await htmlToDocx(docxHtml);
res.contentType(
'application/vnd.openxmlformats-officedocument.wordprocessingml.document');
res.send(buf);
return;
}
if (type === 'pdf') {
const {htmlToPdfBuffer} = require('../utils/ExportPdfNative');
const buf = await htmlToPdfBuffer(bodyHtml);
res.contentType('application/pdf');
res.send(buf);
return;
}
// soffice-only formats (odt, doc) are blocked at the route guard
// when soffice is null; reaching here means the guard is wrong.
res.status(500).send(`Cannot export ${type} without soffice configured`);
return;
} catch (err) {
console.error(
`native ${type} export failed for pad "${padId}":`,
err && (err as Error).stack ? (err as Error).stack : err);
res.status(500).send(`Failed to export pad as ${type}.`);
return;
}
}
// soffice path — write the html export to a file
const randNum = Math.floor(Math.random() * 0xFFFFFFFF);
const srcFile = `${tempDirectory}/etherpad_export_${randNum}.html`;
await fsp_writeFile(srcFile, html);

View file

@ -65,6 +65,11 @@ if (settings.soffice != null) {
exportExtension = 'html';
}
// Office formats with no in-process import path (issue #7538). When soffice
// is null these are rejected explicitly so users see a clear error instead
// of a silent ASCII-only fallback. .docx has a native path via mammoth.
const SOFFICE_ONLY_IMPORT_FORMATS = new Set(['.pdf', '.odt', '.doc', '.rtf']);
const tmpDirectory = os.tmpdir();
/**
@ -130,6 +135,67 @@ const doImport = async (req:any, res:any, padId:string, authorId:string) => {
}
}
// Detect once whether the contentcollector treats h1-h6/code as block
// elements server-side. ep_headings2 v0.2.118+ (after the
// ep_plugin_helpers ccRegisterBlockElements wiring lands) registers
// them; older versions don't. The two preprocessors below are only
// needed when the plugin hook is missing — they're harmful otherwise
// (each adds an extra blank pad line per heading transition).
const ccBlockElems: string[] =
([] as string[]).concat(...(hooks.callAll('ccRegisterBlockElements') || []));
const ccBlockSet = new Set(ccBlockElems.map((t: string) => t.toLowerCase()));
// ep_headings2 registers 'h1' along with the others when its server
// hook is wired (ccRegisterBlockElements). h1 is sufficient as the
// detection probe; the absence of h5/h6 in the set is a quirk of
// ep_headings2 (it only handles h1-h4) and not a sign of a broken
// hook.
const headingsAreBlocks = ccBlockSet.has('h1');
// Native DOCX import (issue #7538): when soffice isn't configured we
// hand .docx files to mammoth, which produces HTML — then we feed that
// through the existing setPadHTML pipeline.
if (settings.soffice == null && fileEnding === '.docx') {
const buf = await fs.readFile(srcFile);
const {docxBufferToHtml} = require('../utils/ImportDocxNative');
const {separateAdjacentHeadingBlocks} = require('../utils/ExportSanitizeHtml');
let nativeHtml: string;
try {
nativeHtml = await docxBufferToHtml(buf);
// When the plugin hook is missing, contentcollector treats h1-h6
// as inline and adjacent headings merge into a single pad line.
// Insert <br> between them as a defensive workaround. Skipped when
// the plugin already registers the tags (otherwise the <br> becomes
// an extra blank line per heading transition).
if (!headingsAreBlocks) {
nativeHtml = separateAdjacentHeadingBlocks(nativeHtml);
}
} catch (err: any) {
logger.warn(`Native DOCX import failed: ${err.stack || err}`);
throw new ImportError('convertFailed');
}
const pad = await padManager.getPad(padId, '\n', authorId);
try {
await importHtml.setPadHTML(pad, nativeHtml, authorId);
} catch (err: any) {
logger.warn(`Error importing native DOCX HTML: ${err.stack || err}`);
throw new ImportError('convertFailed');
}
padManager.unloadPad(padId);
const reloaded = await padManager.getPad(padId, '\n', authorId);
padManager.unloadPad(padId);
await padMessageHandler.updatePadClients(reloaded);
rm(srcFile);
return false;
}
// Without soffice, the legacy office formats (pdf, odt, doc, rtf) have
// no in-process path. Reject explicitly so the user sees a clear error
// instead of a silent ASCII-only fallback.
if (settings.soffice == null && SOFFICE_ONLY_IMPORT_FORMATS.has(fileEnding)) {
logger.warn(`Cannot import ${fileEnding} without soffice configured`);
throw new ImportError('uploadFailed');
}
const destFile = path.join(tmpDirectory, `etherpad_import_${randNum}.${exportExtension}`);
const context = {srcFile, destFile, fileEnding, padId, ImportError};
const importHandledByPlugin = (await hooks.aCallAll('import', context)).some((x:string) => x);
@ -205,7 +271,20 @@ const doImport = async (req:any, res:any, padId:string, authorId:string) => {
if (!directDatabaseAccess) {
if (importHandledByPlugin || useConverter || fileIsHTML) {
try {
await importHtml.setPadHTML(pad, text, authorId);
// Etherpad's HTML export wraps each pad line in `<p>...</p>`
// (or `<h1>`, `<code>`, etc.) and then appends a `<br>` between
// lines. The closing block tag already ends the line for
// contentcollector, so the trailing `<br>` is redundant and
// doubles every blank line on import. Collapse `</block><br>`
// before handing to setPadHTML so HTML round-trips don't drift.
// Only applied to HTML imports (and converted-via-soffice
// outputs, which look the same shape) -- the docx native path
// above doesn't go through here.
const {collapseRedundantBrAfterBlocks} =
require('../utils/ExportSanitizeHtml');
const cleaned = (fileIsHTML || useConverter)
? collapseRedundantBrAfterBlocks(text) : text;
await importHtml.setPadHTML(pad, cleaned, authorId);
} catch (err:any) {
logger.warn(`Error importing, possibly caused by malformed HTML: ${err.stack || err}`);
}

View file

@ -36,9 +36,11 @@ exports.expressCreateServer = (hookName:string, args:ArgsExpressType, cb:Functio
return next();
}
// if soffice is disabled, and this is a format we only support with soffice, output a message
// When soffice is disabled, only block formats with no native path.
// pdf and docx fall through to ExportHandler, which dispatches to
// the in-process converters (issue #7538).
if (exportAvailable() === 'no' &&
['odt', 'pdf', 'doc', 'docx'].indexOf(req.params.type) !== -1) {
['odt', 'doc'].indexOf(req.params.type) !== -1) {
console.error(`Impossible to export pad "${req.params.pad}" in ${req.params.type} format.` +
' There is no converter configured');

View file

@ -0,0 +1,248 @@
'use strict';
import {Parser} from 'htmlparser2';
import {PassThrough} from 'stream';
const PDFDocument = require('pdfkit');
interface InlineState {
bold: boolean;
italic: boolean;
underline: boolean;
strike: boolean;
link?: string;
fontSize?: number;
align?: 'left' | 'center' | 'right' | 'justify';
mono?: boolean;
}
const parseAlign = (style: string | undefined): InlineState['align'] | undefined => {
if (!style) return undefined;
const m = /text-align\s*:\s*(left|center|right|justify)/i.exec(style);
return m ? (m[1].toLowerCase() as InlineState['align']) : undefined;
};
const HEADING_SIZES: Record<string, number> = {
h1: 24, h2: 20, h3: 16, h4: 14, h5: 12, h6: 11,
};
// Tags whose text content must never appear in the rendered PDF (CSS,
// scripts, document metadata). The walker maintains a depth counter so that
// nested elements inside one of these are ignored too.
const SKIP_TAGS = new Set(['head', 'style', 'script', 'title', 'meta', 'link', 'noscript']);
const decodeDataUri = (src: string): Buffer | null => {
const m = /^data:[^;,]+;base64,(.+)$/i.exec(src);
if (!m) return null;
try {
return Buffer.from(m[1], 'base64');
} catch {
return null;
}
};
export const htmlToPdfBuffer = (html: string): Promise<Buffer> =>
new Promise((resolve, reject) => {
// compress:false keeps the content stream uncompressed. Pads are small
// enough that the size cost is negligible, and it lets ops greppable PDFs
// out of the box for accessibility / search-engine indexers that don't
// FlateDecode.
const doc = new PDFDocument({margin: 50, compress: false});
const stream = new PassThrough();
const chunks: Buffer[] = [];
stream.on('data', (c: Buffer) => chunks.push(c));
stream.on('end', () => resolve(Buffer.concat(chunks)));
stream.on('error', reject);
doc.pipe(stream);
const styleStack: InlineState[] = [{
bold: false, italic: false, underline: false, strike: false,
}];
const listType: ('ul' | 'ol' | null)[] = [];
const listIndex: number[] = [];
let pendingNewline = false;
let skipDepth = 0;
const top = () => styleStack[styleStack.length - 1];
const applyFont = () => {
const s = top();
let variant: string;
if (s.mono) {
variant =
s.bold && s.italic ? 'Courier-BoldOblique' :
s.bold ? 'Courier-Bold' :
s.italic ? 'Courier-Oblique' :
'Courier';
} else {
variant =
s.bold && s.italic ? 'Helvetica-BoldOblique' :
s.bold ? 'Helvetica-Bold' :
s.italic ? 'Helvetica-Oblique' :
'Helvetica';
}
doc.font(variant);
doc.fontSize(s.fontSize || 11);
};
// Track whether the current run started with an alignment override so
// we apply `align` exactly once per pdfkit text() call (pdfkit uses the
// align option of the first call in a continued run for the whole line).
let runStartedAligned = false;
const writeText = (raw: string) => {
if (!raw) return;
if (pendingNewline) {
doc.moveDown(0.5);
pendingNewline = false;
}
const s = top();
applyFont();
const opts: any = {continued: true};
if (s.underline) opts.underline = true;
if (s.strike) opts.strike = true;
if (s.link) opts.link = s.link;
if (s.align && !runStartedAligned) {
opts.align = s.align;
runStartedAligned = true;
}
doc.text(raw, opts);
};
// End the current `continued: true` text run. pdfkit's `text('', false)`
// closes the run but does NOT advance the cursor — subsequent text would
// overlay at the same y. Use `breakLine` whenever a true newline is
// intended (br, end-of-block, list items).
const flushLine = () => {
doc.text('', {continued: false});
runStartedAligned = false;
};
const breakLine = () => {
flushLine();
doc.moveDown(1);
};
const parser = new Parser({
onopentag(name, attribs) {
if (SKIP_TAGS.has(name)) skipDepth += 1;
if (skipDepth > 0) {
styleStack.push({...top()});
return;
}
const cur = top();
const next: InlineState = {...cur};
switch (name) {
case 'b': case 'strong': next.bold = true; break;
case 'i': case 'em': next.italic = true; break;
case 'u': next.underline = true; break;
case 's': case 'strike': case 'del': next.strike = true; break;
case 'a': next.link = attribs.href; next.underline = true; break;
case 'code': case 'tt': case 'kbd': case 'samp': {
next.mono = true;
// ep_headings2 uses <code style='text-align:...'> as a block-
// styled "code" line, so read the alignment off the opening
// tag too. parseAlign returns undefined when no text-align
// is set, so this is a no-op for inline <code> usage.
const a = parseAlign(attribs.style);
if (a) next.align = a;
break;
}
case 'pre': {
next.mono = true;
const a = parseAlign(attribs.style);
if (a) next.align = a;
if (!pendingNewline) breakLine();
break;
}
case 'h1': case 'h2': case 'h3': case 'h4': case 'h5': case 'h6': {
next.fontSize = HEADING_SIZES[name];
next.bold = true;
const a = parseAlign(attribs.style);
if (a) next.align = a;
if (!pendingNewline) breakLine();
break;
}
case 'p': case 'div': {
const a = parseAlign(attribs.style);
if (a) next.align = a;
if (!pendingNewline) breakLine();
break;
}
case 'ul': case 'ol':
listType.push(name as 'ul' | 'ol');
listIndex.push(0);
breakLine();
break;
case 'li': {
breakLine();
const t = listType[listType.length - 1] || 'ul';
if (t === 'ol') listIndex[listIndex.length - 1] += 1;
const prefix = t === 'ul'
? '• '
: `${listIndex[listIndex.length - 1]}. `;
const indent = ' '.repeat(Math.max(0, listType.length - 1));
applyFont();
doc.text(`${indent}${prefix}`, {continued: true});
break;
}
case 'br':
breakLine();
break;
case 'img': {
const buf = decodeDataUri(attribs.src || '');
if (buf) {
flushLine();
try { doc.image(buf, {fit: [400, 300]}); } catch { /* skip bad image */ }
}
break;
}
}
styleStack.push(next);
},
ontext(text) {
if (skipDepth > 0) return;
// Collapse consecutive whitespace to a single space, the way an
// HTML renderer would. Without this, literal newlines and tabs in
// pretty-printed source HTML show up as runs of " " in the PDF.
const collapsed = text.replace(/[\s ]+/g, ' ');
if (collapsed === ' ') return; // pure-whitespace runs are dropped
writeText(collapsed);
},
onclosetag(name) {
if (skipDepth > 0) {
if (SKIP_TAGS.has(name)) skipDepth -= 1;
styleStack.pop();
if (styleStack.length === 0) {
styleStack.push({bold: false, italic: false, underline: false, strike: false});
}
return;
}
switch (name) {
case 'h1': case 'h2': case 'h3': case 'h4': case 'h5': case 'h6':
case 'p': case 'div': case 'pre':
breakLine();
pendingNewline = true;
break;
case 'li':
flushLine();
break;
case 'ul': case 'ol':
listType.pop();
listIndex.pop();
doc.moveDown(0.3);
break;
}
styleStack.pop();
if (styleStack.length === 0) {
styleStack.push({bold: false, italic: false, underline: false, strike: false});
}
},
}, {decodeEntities: true, lowerCaseTags: true});
parser.write(html);
parser.end();
flushLine();
doc.end();
});

View file

@ -0,0 +1,215 @@
'use strict';
import {Parser} from 'htmlparser2';
// Pull `<body>...</body>` out of a full HTML document. Etherpad's
// `getPadHTMLDocument()` returns a complete page — `<head>` with a `<style>`
// block, doctype, etc. The legacy LibreOffice path renders that fine, but
// the in-process converters (html-to-docx, our pdfkit walker) treat
// non-body content as renderable, leaking CSS into the output and giving
// blank-line issues from the leading whitespace inside `<body>`. This helper
// extracts the body content and trims surrounding whitespace; if the input
// has no `<body>`, it's returned unchanged so plugin-shaped fragments still
// flow through.
const BODY_RE = /<body[^>]*>([\s\S]*?)<\/body>/i;
export const extractBody = (html: string): string => {
const m = BODY_RE.exec(html);
if (!m) return html;
return m[1].replace(/^[\s ]+/, '').replace(/[\s ]+$/, '');
};
// Drop `<br>` immediately following a closing block tag. Etherpad's
// HTML export writes one `<p>...</p>` per pad line (or `<h1>...</h1>`,
// `<code>...</code>` for the styled ones from ep_align/ep_headings2),
// then appends a `<br>` between lines. The `<br>` is redundant — the
// closing block tag already ends the line — and on import the server
// content collector counts BOTH as line breaks, so every blank line
// between two paragraphs gets duplicated.
const REDUNDANT_BR_RE =
/(<\/(?:p|h[1-6]|div|pre|blockquote|code|ul|ol|li|table|tr|td|th)>)\s*<br\s*\/?>/gi;
export const collapseRedundantBrAfterBlocks = (html: string): string =>
html.replace(REDUNDANT_BR_RE, '$1');
// Insert a `<br>` between adjacent heading-style blocks so etherpad's
// server-side content collector breaks them into separate pad lines.
//
// Background: contentcollector's default `_blockElems` set is just
// `{div, p, pre, li}`. ep_headings2 registers the CLIENT-side
// `aceRegisterBlockElements` for `h1..h4` and `code`, but not the
// SERVER-side `ccRegisterBlockElements`, so on import contentcollector
// treats those tags as inline and merges adjacent ones into a single
// line. This helper fires on the IMPORT path (after mammoth produces
// HTML) to forcibly separate them.
const ADJACENT_HEADING_BLOCKS_RE =
/(<\/(?:h[1-6]|code)>)(\s*<(?:h[1-6]|code|p|pre|div|blockquote|ul|ol)\b)/gi;
export const separateAdjacentHeadingBlocks = (html: string): string =>
html.replace(ADJACENT_HEADING_BLOCKS_RE, '$1<br>$2');
// Convert code/pre/tt/kbd/samp wrappers to plain styled spans (and a
// wrapping <p> when block-styled) so html-to-docx renders them with
// `<w:rFonts w:ascii="Courier New" .../>`. The bare `<code>` tag
// isn't translated to a font change by html-to-docx, AND it has a
// nasty bug where any `<a href>` nested inside `<code>` (or inside a
// styled `<span>`) is silently dropped from the output. Workaround:
// drop the code/pre tag entirely, wrap non-anchor text in monospace
// spans, leave anchors as-is. For block-level usage (e.g.
// ep_headings2's `<code style='text-align:right'>` per-line wrapper)
// we emit a wrapping `<p>` and forward any text-align style.
//
// Run BEFORE `wrapLooseLines` so the resulting `<p>` lands at the
// loose-line boundary instead of getting double-wrapped.
const MONO_TAGS_RE = /<(code|tt|kbd|samp|pre)\b([^>]*)>([\s\S]*?)<\/\1>/gi;
const ANCHOR_RE = /<a\b[^>]*>[\s\S]*?<\/a>/gi;
const STYLE_ATTR_RE = /\bstyle\s*=\s*(['"])([^'"]*)\1/i;
const COURIER_OPEN = '<span style="font-family:\'Courier New\', monospace">';
const COURIER_CLOSE = '</span>';
const wrapNonAnchorSegments = (content: string): string => {
let out = '';
let lastIndex = 0;
let m: RegExpExecArray | null;
ANCHOR_RE.lastIndex = 0;
while ((m = ANCHOR_RE.exec(content)) !== null) {
const before = content.slice(lastIndex, m.index);
if (before) out += `${COURIER_OPEN}${before}${COURIER_CLOSE}`;
out += m[0];
lastIndex = m.index + m[0].length;
}
if (lastIndex < content.length) {
const after = content.slice(lastIndex);
if (after) out += `${COURIER_OPEN}${after}${COURIER_CLOSE}`;
}
return out || `${COURIER_OPEN}${content}${COURIER_CLOSE}`;
};
export const applyMonospaceToCode = (html: string): string =>
html.replace(MONO_TAGS_RE, (_, tag, attrs, content) => {
const styled = wrapNonAnchorSegments(content);
// Block-level treatment for <pre> (always) and <code>/<tt>/etc.
// when the wrapper carries an inline style (ep_headings2 +
// ep_align emit `<code style='text-align:right'>` for each pad
// line). Forward the style to a wrapping `<p>`.
const styleMatch = STYLE_ATTR_RE.exec(attrs);
if (tag.toLowerCase() === 'pre' || styleMatch) {
const styleAttr = styleMatch ? ` style="${styleMatch[2]}"` : '';
return `<p${styleAttr}>${styled}</p>`;
}
return styled;
});
// Drop block elements whose only content is whitespace. Etherpad plugins
// like ep_headings2 emit a heading-styled blank-line block (e.g.
// `<h1 style='text-align:right'></h1>`) after every styled line, which
// turns into an extra empty `<w:p>` in DOCX and an extra blank line in
// PDF. Iterates because removing one empty wrapper can expose another.
//
// Note: `<p></p>` is intentionally NOT in this list — `wrapLooseLines`
// uses empty `<p>` markers to encode blank-line gaps for round-trip
// fidelity through html-to-docx.
const EMPTY_BLOCK_RE = /<(h[1-6]|code|pre|div|blockquote)\b[^>]*>\s*<\/\1>/gi;
export const dropEmptyBlocks = (html: string): string => {
let prev: string;
let cur = html;
do {
prev = cur;
cur = cur.replace(EMPTY_BLOCK_RE, '');
} while (cur !== prev);
return cur;
};
// Wrap loose text + inline content in `<p>` blocks so html-to-docx renders
// `<br>` as a soft line break (`<w:br/>`) instead of a paragraph break
// (`<w:p>`). Etherpad's HTML export uses bare `<br>` for every line and
// `<br><br>` for blank lines, so without this DOCX exports get one Word
// paragraph per line and two empty paragraphs for every blank line.
//
// Strategy: capture `<br>` separators of length >= 2 (paragraph separators)
// AND remember how many `<br>`s each separator contains, so blank-line
// gaps survive the round-trip. For N consecutive `<br>`s, emit one
// closing-then-opening paragraph break PLUS (N - 2) empty `<p></p>`
// markers (each empty paragraph = one blank pad line).
const BLOCK_HEAD_RE = /^<(?:p|h[1-6]|ul|ol|table|blockquote|pre|div)[\s>/]/i;
// Anchored so the inner `\s*` can't overlap with surrounding whitespace and
// trigger exponential backtracking. Matches `<br>` followed by at least one
// more `<br>` (with optional whitespace between).
const BR_PARA_RE = /<br\s*\/?>(?:\s*<br\s*\/?>)+/gi;
const TRAILING_BR_RE = /(?:<br\s*\/?>\s*)+$/i;
const BR_COUNT_RE = /<br/gi;
export const wrapLooseLines = (html: string): string => {
// split() with a capturing group keeps the separators in the result, so
// parts[i] alternates between content (even i) and br-run separator
// (odd i).
const parts = html.split(/(<br\s*\/?>(?:\s*<br\s*\/?>)+)/gi);
const out: string[] = [];
for (let i = 0; i < parts.length; i++) {
if (i % 2 === 0) {
const c = parts[i].replace(TRAILING_BR_RE, '').trim();
if (!c) continue;
out.push(BLOCK_HEAD_RE.test(c) ? c : `<p>${c}</p>`);
} else {
// Separator of N >= 2 <br>s. The first <br> is the paragraph
// boundary; the remaining (N - 1) each represent one blank pad
// line, emitted as an empty <p></p>.
const n = (parts[i].match(BR_COUNT_RE) || []).length;
for (let k = 0; k < n - 1; k++) out.push('<p></p>');
}
}
return out.join('');
};
const isLocalSrc = (src: string): boolean => {
if (!src) return true;
if (src.startsWith('data:')) return true;
if (src.startsWith('//')) return false;
if (/^[a-z][a-z0-9+.-]*:/i.test(src)) return false;
return true;
};
const escapeAttr = (s: string): string =>
s.replace(/&/g, '&amp;').replace(/"/g, '&quot;').replace(/</g, '&lt;');
const escapeText = (s: string): string =>
s.replace(/&/g, '&amp;').replace(/</g, '&lt;').replace(/>/g, '&gt;');
const VOID_TAGS = new Set([
'area', 'base', 'br', 'col', 'embed', 'hr', 'img', 'input',
'link', 'meta', 'source', 'track', 'wbr',
]);
export const stripRemoteImages = (html: string): string => {
let out = '';
const parser = new Parser({
onopentag(name, attribs) {
if (name === 'img') {
const src = attribs.src || '';
if (isLocalSrc(src)) {
let tag = '<img';
for (const [k, v] of Object.entries(attribs)) {
tag += ` ${k}="${escapeAttr(v)}"`;
}
tag += '>';
out += tag;
} else {
out += escapeText(attribs.alt || '');
}
return;
}
let tag = `<${name}`;
for (const [k, v] of Object.entries(attribs)) {
tag += ` ${k}="${escapeAttr(v)}"`;
}
tag += '>';
out += tag;
},
ontext(text) {
out += text;
},
onclosetag(name) {
if (VOID_TAGS.has(name)) return;
out += `</${name}>`;
},
}, {decodeEntities: false, lowerCaseTags: true});
parser.write(html);
parser.end();
return out;
};

View file

@ -0,0 +1,83 @@
'use strict';
const mammoth = require('mammoth');
const JSZip = require('jszip');
// mammoth strips paragraph alignment (<w:jc>) when it converts a docx to
// HTML; it has no equivalent style-mapping for justification. To keep
// alignment through the round-trip we walk the docx's document.xml
// directly, pull the `w:val` from each `<w:p>`'s `<w:jc>`, and inject a
// matching `style="text-align:..."` onto the corresponding block element
// in mammoth's output. Match is by document order: the Nth `<p>` /
// `<h1>...<h6>` in mammoth's output corresponds to the Nth `<w:p>` in
// the docx.
const PARA_RE = /<w:p\b[^>]*>([\s\S]*?)<\/w:p>/g;
const JC_RE = /<w:jc\s+w:val=["']([^"']+)["']/;
// Word's `w:jc` accepts more values than CSS text-align; map the ones
// we want to surface and skip the rest.
const JC_TO_CSS: Record<string, string> = {
left: 'left',
start: 'left',
center: 'center',
right: 'right',
end: 'right',
both: 'justify',
justify: 'justify',
distribute: 'justify',
};
const extractAlignmentMap = async (buffer: Buffer): Promise<Array<string|null>> => {
const aligns: Array<string|null> = [];
try {
const zip = await JSZip.loadAsync(buffer);
const file = zip.file('word/document.xml');
if (!file) return aligns;
const xml: string = await file.async('text');
let m: RegExpExecArray | null;
while ((m = PARA_RE.exec(xml)) !== null) {
const jcMatch = JC_RE.exec(m[1]);
const css = jcMatch ? JC_TO_CSS[jcMatch[1].toLowerCase()] : null;
aligns.push(css || null);
}
} catch {
// Best-effort — if the docx structure is anything other than a
// standard document.xml, fall back to no alignment.
}
return aligns;
};
const applyAlignmentToHtml = (html: string, aligns: Array<string|null>): string => {
if (aligns.length === 0 || !aligns.some((a) => a && a !== 'left')) return html;
let i = 0;
return html.replace(/<(p|h[1-6])(\b[^>]*)>/gi, (whole: string, tag: string, attrs: string) => {
const align = aligns[i++];
if (!align || align === 'left') return whole;
if (/\bstyle\s*=/.test(attrs)) {
return `<${tag}${attrs.replace(
/\bstyle\s*=\s*(['"])([^'"]*)\1/i,
(_full: string, q: string, val: string) => `style=${q}${val}; text-align:${align}${q}`)}>`;
}
return `<${tag}${attrs} style="text-align:${align}">`;
});
};
export const docxBufferToHtml = async (buffer: Buffer): Promise<string> => {
const aligns = await extractAlignmentMap(buffer);
const result = await mammoth.convertToHtml(
{buffer},
{
// Preserve empty paragraphs so blank pad lines survive a
// round-trip. mammoth defaults to true and drops them, which
// collapses blank lines in the middle of a pad's content.
ignoreEmptyParagraphs: false,
convertImage: mammoth.images.imgElement(async (image: any) => {
const buf: Buffer = await image.read();
const contentType = image.contentType || 'application/octet-stream';
return {src: `data:${contentType};base64,${buf.toString('base64')}`};
}),
},
);
let html: string = result.value || '';
html = applyAlignmentToHtml(html, aligns);
return html;
};