The Challenge
The platform’s publishing operation depends on converting thousands of legacy Gencode files—the proprietary format municipal legal editors have used for decades—into clean, structured XHTML/XML that its new oXygen and TopLeaf environment can render and publish. And files had, on average, 800+ pages.
The client’s internal Gencode-to-HTML5 converter existed as a proof of concept, not as a production system. Every converted publication needed significant manual intervention:
A 350-page publication required roughly 2.5 hours of manual editor cleanup after every conversion.
Output XML carried wrong section data types, missing metadata tags, malformed footnotes, broken cross-references, and invalid table markup.
The output folder structure was non-standard, causing downstream failures in the publishing pipeline.
There was no support for Partial Chapter Replacements (PCR)—a routine legal publishing operation where only part of a chapter is updated.
Editors used the converter as a starting point, never as a finished result.
The Approach
Rather than rebuilding the converter from scratch, the team mapped every conversion failure category and prioritized it by frequency and severity. Three inputs shaped the roadmap:
The client’s itemized enhancement requests, compiled from real editor feedback across multiple publications.
Testing feedback spreadsheets produced by the GovTech company’s editors after each release, detailing what still needed correction.
Sprint reviews and team syncs every two weeks that kept stakeholders directly in the feedback loop.
The engagement ran on a two-week sprint cadence. Each sprint closed with a demo and a new versioned release distributed to the client’s team for testing—creating a tight improvement loop between engineering and editorial.
The Solution
We worked on:
XML Structure and Section Handling: Corrected section data-type assignment across every converted file, data-break-before attribute applied to root sections for proper pagination, and standardized content block wrapping in preliminary sections.
Cross-References and Linking: Automated detection and linking of internal and cross-file references, extended cross-reference resolution to the section level with consistent, schema-compliant markup.
Tables, Footnotes, and Indexes: Fixed table formatting, footnote markup placement, and index markup across all files, correctly distinguishing footnote markers from index asterisks.
Metadata and Publishing Schema: Aligned metadata fields to the new publishing schema, including automatic legislation tagging, additional metadata fields, and cleanup of legacy class and preset attributes.
Output, Files, and Assets: Enforced the required output folder structure and corrected asset paths, eliminating the downstream failures that had been breaking the publishing pipeline.
Partial Chapter Replacements: Designed and built the full PCR workflow from scratch—Insert Instructions and List of Effective Pages generation, page-pairing logic for prefolio ranges, first-publication handling, and supplement footer rendering—automating a process the tool couldn't previously handle at all.

The Outcomes
By the end of Sprint 9, the converter had moved from a tool editors used cautiously to a reliable conversion engine producing publication-ready XML across document types and complexity levels.
11 production releases shipped across nine two-week sprints.
540+ tracked issues resolved, covering structural XML errors, metadata gaps, output pipeline failures, and complex editorial edge cases.
Partial Chapter Replacement workflow fully automated—a supplemental process the converter could not previously handle at all.
An editor feedback loop established, giving the client’s team direct influence over each release’s priority order.
Our partner requested an extension to expand the engagement beyond the original scope, a direct signal that the embedded squad (under a Team as a Service model) delivered value well beyond what was expected.


