Legacy system modernization: turning a manual XML cleanup process into an automated conversion engine

A leading GovTech platform serving local governments across the U.S. was modernizing its legal publishing stack, replacing Gencode, Word, and XPP with oXygen, XHTML, and TopLeaf. Its internal conversion tool produced XML that required hours of manual correction before anything could be published. We turned that tool into a production-ready conversion engine. Over nine two-week sprints, an embedded Arionkoder squad shipped 11 production releases, resolved 540+ tracked issues, and built a complete Partial Chapter Replacement workflow the tool could not previously handle at all.

1.4x

file processing speed

540+

tracked issues resolved

11

production releases shipped in under four months

The Challenge

The platform’s publishing operation depends on converting thousands of legacy Gencode files—the proprietary format municipal legal editors have used for decades—into clean, structured XHTML/XML that its new oXygen and TopLeaf environment can render and publish. And files had, on average, 800+ pages.

The client’s internal Gencode-to-HTML5 converter existed as a proof of concept, not as a production system. Every converted publication needed significant manual intervention:

  • A 350-page publication required roughly 2.5 hours of manual editor cleanup after every conversion.

  • Output XML carried wrong section data types, missing metadata tags, malformed footnotes, broken cross-references, and invalid table markup.

  • The output folder structure was non-standard, causing downstream failures in the publishing pipeline.

  • There was no support for Partial Chapter Replacements (PCR)—a routine legal publishing operation where only part of a chapter is updated.

  • Editors used the converter as a starting point, never as a finished result.

The Approach

Rather than rebuilding the converter from scratch, the team mapped every conversion failure category and prioritized it by frequency and severity. Three inputs shaped the roadmap:

  • The client’s itemized enhancement requests, compiled from real editor feedback across multiple publications.

  • Testing feedback spreadsheets produced by the GovTech company’s editors after each release, detailing what still needed correction.

  • Sprint reviews and team syncs every two weeks that kept stakeholders directly in the feedback loop.

The engagement ran on a two-week sprint cadence. Each sprint closed with a demo and a new versioned release distributed to the client’s team for testing—creating a tight improvement loop between engineering and editorial.

The Solution

We worked on:

XML Structure and Section Handling: Corrected section data-type assignment across every converted file, data-break-before attribute applied to root sections for proper pagination, and standardized content block wrapping in preliminary sections.

Cross-References and Linking: Automated detection and linking of internal and cross-file references, extended cross-reference resolution to the section level with consistent, schema-compliant markup.

Tables, Footnotes, and Indexes: Fixed table formatting, footnote markup placement, and index markup across all files, correctly distinguishing footnote markers from index asterisks.

Metadata and Publishing Schema: Aligned metadata fields to the new publishing schema, including automatic legislation tagging, additional metadata fields, and cleanup of legacy class and preset attributes. 

Output, Files, and Assets: Enforced the required output folder structure and corrected asset paths, eliminating the downstream failures that had been breaking the publishing pipeline.

Partial Chapter Replacements: Designed and built the full PCR workflow from scratch—Insert Instructions and List of Effective Pages generation, page-pairing logic for prefolio ranges, first-publication handling, and supplement footer rendering—automating a process the tool couldn't previously handle at all.

The Outcomes

By the end of Sprint 9, the converter had moved from a tool editors used cautiously to a reliable conversion engine producing publication-ready XML across document types and complexity levels.

  • 11 production releases shipped across nine two-week sprints.

  • 540+ tracked issues resolved, covering structural XML errors, metadata gaps, output pipeline failures, and complex editorial edge cases.

  • Partial Chapter Replacement workflow fully automated—a supplemental process the converter could not previously handle at all.

  • An editor feedback loop established, giving the client’s team direct influence over each release’s priority order.

  • Our partner requested an extension to expand the engagement beyond the original scope, a direct signal that the embedded squad (under a Team as a Service model) delivered value well beyond what was expected.

file processing speed

tracked issues resolved

production releases shipped in under four months

1.

2.

3.

4.

5.

Get Started

Ready to make AI useful?

Turning bold ambition into lasting impact starts with a conversation.

Turning bold ambition into lasting impact starts with a conversation.

© 2025 Arionkoder. All rights reserved.

© 2025 Arionkoder. All rights reserved.