commit 2cb04c77a9310e566ca94a40e7c84f8f19f108fe Author: Ken Schaefer Date: Wed Jul 15 16:55:37 2026 -0500 Initial commit containing architetural scaffolding and mvp diff --git a/.thoth/00-project-state.md b/.thoth/00-project-state.md new file mode 100644 index 0000000..c49d9dd --- /dev/null +++ b/.thoth/00-project-state.md @@ -0,0 +1,31 @@ +Current Version + +0.0.1 + +Current Focus + +ChatGPT Capture Connector + +Completed + +✓ Reference Architecture + +✓ ADR-001 + +✓ ADR-002 + +✓ ADR-003 + +✓ Metadata Generator + +In Progress + +• Capture Service + +Blocked + +• Browser Extension + +Next Milestone + +Capture MVP \ No newline at end of file diff --git a/.thoth/10-architecture.md b/.thoth/10-architecture.md new file mode 100644 index 0000000..6d4a691 --- /dev/null +++ b/.thoth/10-architecture.md @@ -0,0 +1,341 @@ +# Project Thoth Architecture + +**Version:** 1.0 +**Status:** Living Document + +--- + +# Purpose + +This document describes the canonical architecture of Project Thoth. + +It is intended to answer a single question: + +> **How does information flow through the Project Thoth ecosystem?** + +This document describes architecture, not implementation. + +Current implementation status is maintained separately in `.thoth/00-project-state.md`. + +--- + +# Architectural Principles + +Project Thoth is built upon the following principles: + +- Source artifacts are permanent. +- Generated artifacts are disposable. +- Every transformation is deterministic. +- AI providers are interchangeable. +- Markdown is the canonical document format. +- Knowledge is owned by the archive, not the AI. +- Processing stages should be loosely coupled. +- Every stage should be independently testable. + +--- + +# High-Level Architecture + +``` + External Sources + │ + ▼ + Capture Connectors + │ + ▼ + Canonical Source Archive + │ + ▼ + Processing Pipeline + │ + ┌──────────────┼──────────────┐ + ▼ ▼ ▼ + Metadata Manifest Harvests + Generation Generation Generation + │ │ │ + └──────────────┼──────────────┘ + ▼ + Knowledge Repository + │ + ▼ + Search / Retrieval Layer + │ + ▼ + AI Consumers +``` + +--- + +# Layers + +## 1. External Sources + +Project Thoth captures knowledge from external systems. + +Examples include: + +- ChatGPT +- Claude +- Gemini +- Microsoft Copilot +- Open WebUI +- Local LLMs +- Documentation systems +- Markdown repositories +- Future knowledge sources + +These systems are considered transient. + +They are never considered the canonical repository. + +--- + +## 2. Capture Connectors + +Capture Connectors preserve source material. + +Responsibilities include: + +- Capture conversations +- Preserve ordering +- Preserve formatting +- Preserve attachments +- Produce canonical source documents + +Capture Connectors do **not**: + +- Summarize +- Interpret +- Classify +- Generate metadata +- Perform AI reasoning + +They are intentionally deterministic. + +--- + +## 3. Canonical Source Archive + +The archive contains the preserved source material. + +Typical artifacts include: + +- conversation.md +- source assets +- attachments + +The archive is immutable. + +Once source material has been captured it should not be modified. + +If processing changes in the future, artifacts are regenerated from the archived source. + +--- + +## 4. Processing Pipeline + +Processors transform source artifacts into derived artifacts. + +Typical processors include: + +- Metadata generation +- Manifest generation +- Harvest generation +- Search indexing +- Embedding generation + +Processors never modify source artifacts. + +They only generate derived artifacts. + +--- + +## 5. Knowledge Repository + +The repository contains: + +- Source artifacts +- Derived artifacts +- Metadata +- Relationships + +This becomes the canonical knowledge base for Project Thoth. + +--- + +## 6. Consumers + +Consumers interact with the repository. + +Examples: + +- Local LLMs +- Search engines +- Knowledge assistants +- Documentation systems +- Human readers + +Consumers should never depend on the original AI platform. + +--- + +# Capture Connector Architecture + +Every Capture Connector implements the same logical pipeline. + +``` +Optional Preprocessing + │ + ▼ +Conversation Discovery + │ + ▼ +Intermediate Representation + │ + ▼ +Content Transformation + │ + ▼ +Markdown Serialization + │ + ▼ +conversation.md +``` + +This architecture is defined by ADR-002. + +Only the Discovery stage is expected to be platform-specific. + +--- + +# Processing Philosophy + +Project Thoth separates preservation from interpretation. + +``` +Capture Once + +↓ + +Preserve Forever + +↓ + +Process Many Times +``` + +This allows improvements to processors without requiring the original source to be recaptured. + +--- + +# Canonical Artifacts + +Project Thoth distinguishes between canonical artifacts and generated artifacts. + +## Canonical + +Examples: + +- conversation.md +- Source attachments +- Images +- Audio +- Video + +Canonical artifacts are permanent. + +--- + +## Generated + +Examples: + +- Metadata +- Manifests +- Harvests +- Embeddings +- Search indexes + +Generated artifacts may be deleted and regenerated at any time. + +--- + +# Repository Structure + +A typical Project Thoth repository contains: + +``` +.thoth/ + Project self-description + +applications/ + Executable applications + +processors/ + Artifact processors + +specifications/ + Canonical document specifications + +docs/ + Architecture and ADRs + +codex/ + Engineering work orders + +archive/ + Captured source material + +generated/ + Derived artifacts +``` + +Implementation details may evolve while preserving the architectural boundaries described in this document. + +--- + +# Architectural Boundaries + +Each layer has a single responsibility. + +| Layer | Responsibility | +|---------|----------------| +| Capture Connectors | Preserve source material | +| Archive | Store canonical source | +| Processors | Generate derived artifacts | +| Repository | Organize knowledge | +| Consumers | Read knowledge | + +No layer should assume the responsibilities of another. + +--- + +# Design Goals + +The architecture is intended to achieve: + +- Long-term preservation +- Platform independence +- Deterministic processing +- Reproducibility +- Testability +- Portability +- AI interoperability + +--- + +# Relationship to Other Documents + +This document defines **how Project Thoth is organized**. + +Related documents include: + +- **00-project-state.md** — Current implementation status. +- **20-philosophy.md** — Guiding principles. +- **30-glossary.md** — Canonical terminology. +- **40-coding-standards.md** — Development standards. +- **ADR-001** — Project structure. +- **ADR-002** — Capture connector pipeline. +- **ADR-003** — Capture connector architecture. + +Together these documents form the canonical self-description of the project. \ No newline at end of file diff --git a/.thoth/20-philosophy.md b/.thoth/20-philosophy.md new file mode 100644 index 0000000..d650a1d --- /dev/null +++ b/.thoth/20-philosophy.md @@ -0,0 +1,321 @@ +# Project Thoth Philosophy + +**Version:** 1.0 +**Status:** Living Document + +--- + +# Purpose + +Project Thoth is founded on a simple belief: + +> **Knowledge should outlive the systems that create it.** + +Artificial intelligence systems, software platforms, file formats, and vendors will continue to evolve. + +Human knowledge should not depend on any of them. + +Project Thoth exists to preserve human reasoning in a portable, deterministic, and vendor-independent form. + +--- + +# Core Philosophy + +Project Thoth is not an AI project. + +It is a knowledge preservation project that happens to use AI. + +Artificial intelligence is a consumer of knowledge, not its owner. + +The archive is the product. + +Everything else is replaceable. + +--- + +# Guiding Principles + +## 1. The Archive is Canonical + +The archive is the source of truth. + +Not ChatGPT. + +Not Claude. + +Not Gemini. + +Not a vector database. + +Not an embedding model. + +Every generated artifact should be reproducible from the archive. + +--- + +## 2. Source Artifacts are Permanent + +Original captured material should never be destroyed. + +Examples include: + +- Conversations +- Images +- Audio +- Video +- Attachments + +Source artifacts represent historical evidence. + +They cannot be recreated once lost. + +--- + +## 3. Generated Artifacts are Disposable + +Metadata. + +Embeddings. + +Search indexes. + +Harvests. + +Summaries. + +These exist only to improve access to the archive. + +They should always be considered disposable. + +If a better algorithm is developed tomorrow, they should simply be regenerated. + +--- + +## 4. Capture Once. Process Forever. + +Knowledge capture should occur only once. + +Processing should occur as many times as necessary. + +New processors should never require recollecting historical conversations. + +--- + +## 5. Determinism Over Magic + +Project Thoth prefers deterministic transformations over opaque intelligence. + +The same input should produce the same output. + +When AI reasoning is introduced, it should be clearly separated from preservation. + +--- + +## 6. AI Providers are Interchangeable + +No AI vendor should become part of the archive. + +ChatGPT. + +Claude. + +Gemini. + +Copilot. + +Open WebUI. + +Future models. + +These are simply sources of knowledge. + +The archive must remain usable regardless of which systems disappear. + +--- + +## 7. Markdown is the Universal Language + +Project Thoth prefers open, human-readable formats. + +Markdown provides: + +- longevity +- portability +- version control +- interoperability +- readability + +Whenever practical, proprietary formats should be transformed into Markdown. + +--- + +## 8. Knowledge Should be Human-Readable + +Every canonical artifact should be understandable without specialized software. + +A person should be able to browse the archive with nothing more than a text editor. + +Software enhances the archive. + +It should never become a prerequisite for understanding it. + +--- + +## 9. Preserve Before Improving + +Improvement can always happen later. + +Preservation cannot. + +When faced with a choice between: + +- perfect processing +- complete preservation + +Project Thoth chooses preservation. + +Incomplete but preserved information is preferable to perfectly formatted information that has silently lost content. + +--- + +## 10. Separation of Responsibilities + +Every component should have one responsibility. + +Capture preserves. + +Processors transform. + +Indexes organize. + +AI reasons. + +Mixing these concerns creates unnecessary complexity. + +--- + +## 11. Reproducibility Matters + +Every generated artifact should be reproducible. + +A future developer should be able to regenerate: + +- metadata +- manifests +- harvests +- search indexes +- embeddings + +from the canonical archive without requiring access to the original AI platform. + +--- + +## 12. Preserve Context, Not Just Content + +Facts without context lose meaning. + +Whenever possible, preserve: + +- ordering +- timestamps +- relationships +- authorship +- source platform +- surrounding discussion + +Knowledge is more than isolated documents. + +It is the reasoning that connects them. + +--- + +## 13. The Project Describes Itself + +Project Thoth repositories should be self-describing. + +A developer. + +An AI assistant. + +An automated process. + +All should be able to understand the project by reading its canonical documents. + +Knowledge about the project belongs inside the project. + +--- + +## 14. Architecture Should Outlive Implementation + +Applications will be rewritten. + +Frameworks will change. + +Programming languages will evolve. + +The architecture should remain recognizable. + +Implementation serves architecture. + +Architecture serves philosophy. + +--- + +# What Project Thoth is Not + +Project Thoth is not: + +- another note-taking application +- another vector database +- another chatbot +- another RAG framework +- another AI assistant + +These technologies may become components of a Project Thoth ecosystem. + +They are not its purpose. + +--- + +# Long-Term Vision + +Project Thoth seeks to become a durable knowledge architecture. + +A place where: + +- conversations survive software vendors +- ideas survive AI models +- archives survive databases +- knowledge survives technology shifts + +The objective is not simply to remember information. + +It is to preserve reasoning. + +--- + +# Decision Filter + +When making architectural decisions, ask: + +1. Does this strengthen the archive? + +2. Is the result deterministic? + +3. Can it be regenerated? + +4. Does it reduce vendor dependence? + +5. Will this still make sense in ten years? + +If the answer to these questions is "yes," the decision is likely aligned with the philosophy of Project Thoth. + +--- + +# Final Principle + +Technology changes. + +Knowledge endures. + +Project Thoth exists to ensure that the second is never held hostage by the first. \ No newline at end of file diff --git a/.thoth/30-glossary.md b/.thoth/30-glossary.md new file mode 100644 index 0000000..9e8d11a --- /dev/null +++ b/.thoth/30-glossary.md @@ -0,0 +1,461 @@ +# Project Thoth Glossary + +**Version:** 1.0 +**Status:** Canonical Terminology + +--- + +# Purpose + +This glossary defines the canonical vocabulary used throughout Project Thoth. + +The purpose of this document is to ensure that humans, AI assistants, and automated tooling use the same terminology consistently. + +Where possible, each concept has exactly one preferred term. + +--- + +# A + +## Archive + +The permanent collection of canonical source artifacts captured by Project Thoth. + +The archive is the authoritative source of knowledge for the project. + +Examples include: + +- conversations +- images +- attachments +- audio +- video + +The archive is immutable. + +--- + +## Artifact + +Any file produced or managed by Project Thoth. + +Artifacts fall into two categories: + +- Canonical Artifacts +- Generated Artifacts + +--- + +# C + +## Canonical Artifact + +An artifact considered the permanent source of truth. + +Examples: + +- `conversation.md` +- Original attachments +- Images +- Audio +- Video + +Canonical artifacts are preserved indefinitely. + +--- + +## Capture + +The act of preserving information from an external source. + +Capture is deterministic. + +Capture never interprets the information being preserved. + +--- + +## Capture Connector + +A platform-specific adapter responsible for preserving source material. + +Responsibilities include: + +- Discovering conversations +- Preserving formatting +- Preserving ordering +- Producing canonical source artifacts + +Capture Connectors never perform semantic analysis or AI reasoning. + +Examples: + +- ChatGPT Connector +- Claude Connector +- Gemini Connector +- Open WebUI Connector + +--- + +## Capture Pipeline + +The sequence of stages that transforms external content into canonical Project Thoth artifacts. + +Defined in ADR-002. + +``` +Discovery + ↓ +Intermediate Representation + ↓ +Transformation + ↓ +Serialization +``` + +--- + +## Collection + +A logical grouping of related canonical artifacts. + +Collections may represent: + +- projects +- research topics +- books +- clients +- domains + +Collections organize knowledge without altering source artifacts. + +--- + +## Conversation + +A chronological exchange between one or more participants. + +A conversation is preserved as a canonical artifact. + +--- + +## Conversation Turn + +A single exchange within a conversation. + +A turn represents exactly one logical contribution from a participant. + +A turn may contain: + +- paragraphs +- lists +- tables +- images +- code blocks +- attachments + +A turn is never divided into multiple Project Thoth messages. + +--- + +# D + +## Derived Artifact + +An artifact generated from canonical source material. + +Derived artifacts may always be regenerated. + +Examples include: + +- metadata +- manifests +- harvests +- embeddings +- search indexes + +--- + +## Discovery + +The process of locating logical content within an external system. + +Discovery is platform-specific. + +Discovery does not perform formatting or interpretation. + +--- + +# G + +## Generated Artifact + +See **Derived Artifact**. + +--- + +# H + +## Harvest + +A synthesized knowledge artifact produced from one or more canonical sources. + +Harvests represent interpretation rather than preservation. + +Harvests are always generated. + +--- + +# I + +## Intermediate Representation (IR) + +The platform-neutral model produced after discovery. + +The IR separates platform-specific extraction from platform-independent processing. + +The IR is not normally stored permanently. + +--- + +# K + +## Knowledge Repository + +The organized collection of canonical and generated artifacts managed by Project Thoth. + +The repository represents the project's complete preserved knowledge. + +--- + +# M + +## Manifest + +A generated document describing a collection of artifacts. + +A manifest typically records: + +- included files +- relationships +- metadata +- processing status + +Manifests are generated and may be regenerated at any time. + +--- + +## Markdown + +The canonical document format used by Project Thoth. + +Markdown is preferred because it is: + +- human-readable +- portable +- version-control friendly +- widely supported + +--- + +## Metadata + +Structured information describing an artifact. + +Examples include: + +- title +- author +- capture date +- source platform +- tags +- relationships + +Metadata is generated from canonical artifacts. + +--- + +# P + +## Philosophy + +The enduring principles that guide architectural decisions. + +The philosophy explains *why* Project Thoth is designed the way it is. + +--- + +## Processor + +A platform-independent component that transforms canonical artifacts into generated artifacts. + +Examples include: + +- Metadata Generator +- Manifest Generator +- Harvest Generator + +Processors never modify canonical artifacts. + +--- + +## Project Self-Description + +The collection of documents contained within the `.thoth` directory. + +These documents allow humans and AI systems to understand the project without relying on external knowledge. + +--- + +# R + +## Repository + +The complete Project Thoth project, including: + +- applications +- archive +- processors +- documentation +- specifications +- work orders +- canonical project description + +--- + +# S + +## Serialization + +The process of writing transformed content into a canonical Project Thoth document. + +Examples include: + +- `conversation.md` +- metadata files +- manifests + +--- + +## Source Artifact + +A canonical artifact captured directly from an external source. + +Source artifacts are never regenerated. + +--- + +## Source Platform + +The external system from which information was captured. + +Examples: + +- ChatGPT +- Claude +- Gemini +- Microsoft Copilot +- Open WebUI + +--- + +## Specification + +A document defining the canonical structure of Project Thoth artifacts. + +Specifications describe formats rather than implementations. + +--- + +# T + +## Transformation + +The process of converting source content into canonical Project Thoth formats. + +Examples include: + +- HTML → Markdown +- DOM → Intermediate Representation + +Transformation does not alter meaning. + +--- + +## Turn Renderer + +A component responsible for converting one Conversation Turn into canonical Markdown. + +A Turn Renderer never discovers conversation turns. + +--- + +# V + +## Vault + +A user-managed knowledge repository containing canonical Project Thoth artifacts. + +The vault is tool-independent. + +Applications such as Obsidian, VS Code, or future tools may operate on the same vault. + +The vault is not defined by any particular application. + +--- + +# W + +## Work Order + +A structured engineering document describing a discrete unit of implementation work. + +A Work Order contains: + +- context +- objectives +- constraints +- acceptance criteria +- architectural guidance + +Work Orders are intended to be executable by either human developers or AI assistants. + +Work Orders are stored under the `codex/` directory. + +--- + +# Canonical Terminology + +The following preferred terms should be used consistently throughout Project Thoth. + +| Preferred | Avoid | +|-----------|-------| +| Archive | Database (when referring to preserved knowledge) | +| Capture Connector | Scraper | +| Conversation Turn | Message Block | +| Derived Artifact | Output File | +| Discovery | DOM Parsing | +| Intermediate Representation | Temporary Object | +| Markdown Serialization | Export | +| Processor | Script | +| Source Artifact | Raw File | +| Vault | Obsidian Vault | + +--- + +# Naming Principles + +When introducing new terminology: + +1. Prefer nouns over verbs. +2. Use one canonical term for each concept. +3. Avoid vendor-specific language. +4. Prefer platform-neutral terminology. +5. Preserve consistency over novelty. + +--- + +# Final Principle + +A shared vocabulary is part of the project's architecture. + +When every contributor—human or AI—uses the same words to describe the same concepts, communication becomes more precise, architectural decisions become easier to understand, and knowledge becomes more durable. \ No newline at end of file diff --git a/.thoth/40-coding-standards.md b/.thoth/40-coding-standards.md new file mode 100644 index 0000000..bea3fe2 --- /dev/null +++ b/.thoth/40-coding-standards.md @@ -0,0 +1,429 @@ +# Project Thoth Coding Standards + +**Version:** 1.0 +**Status:** Canonical Development Standard + +--- + +# Purpose + +This document defines the engineering standards used throughout Project Thoth. + +The objective is not merely to produce working software, but to produce software that remains understandable, maintainable, and portable over time. + +These standards apply equally to: + +- Human developers +- AI-assisted development +- Autonomous development agents + +When implementation choices conflict with these standards, architectural decisions (ADRs) take precedence. + +--- + +# Core Principles + +## 1. Architecture Before Code + +Every significant implementation should begin with architecture. + +If a design decision is expected to persist beyond a single implementation, document it in an Architecture Decision Record (ADR) before coding. + +Code should implement architecture—not define it. + +--- + +## 2. Prefer Refactoring Over Patching + +When repeated fixes indicate a flawed design, refactor rather than accumulate patches. + +Signs that refactoring is warranted include: + +- duplicate logic +- increasing conditional complexity +- repeated bug fixes in the same area +- unclear module responsibilities +- growing technical debt + +The goal is long-term simplicity, not short-term completion. + +--- + +## 3. Single Responsibility + +Every module should have one clear responsibility. + +Examples: + +Good: + +- Conversation Discovery +- HTML-to-Markdown Conversion +- Markdown Serialization +- Metadata Generation + +Poor: + +- ExtractConversationAndGenerateMarkdownAndDownload() + +If a component's name requires multiple conjunctions ("and", "or"), it likely has more than one responsibility. + +--- + +## 4. Separation of Concerns + +Capture, transformation, processing, and presentation are independent concerns. + +Each should exist in its own layer. + +No layer should assume the responsibilities of another. + +--- + +## 5. Deterministic Behavior + +Given the same input, Project Thoth components should produce the same output. + +Randomness, hidden state, and side effects should be avoided unless explicitly required. + +--- + +# Repository Organization + +The repository should separate concerns by purpose rather than technology. + +Typical structure: + +```text +.thoth/ +applications/ +archive/ +codex/ +docs/ +processors/ +specifications/ +``` + +Each top-level directory should represent a distinct architectural concern. + +--- + +# Module Design + +Modules should be: + +- cohesive +- loosely coupled +- independently testable + +Modules should expose small, well-defined interfaces. + +Avoid exposing internal implementation details. + +--- + +# Function Design + +Functions should: + +- perform one task +- have descriptive names +- avoid unnecessary side effects +- return predictable results + +Prefer composition over deeply nested logic. + +--- + +# Error Handling + +Errors should be: + +- explicit +- actionable +- recoverable when practical + +Never silently discard data. + +When partial processing is possible: + +- preserve available information +- report the failure +- continue processing where safe + +Preservation takes precedence over perfection. + +--- + +# Logging + +Logs should assist debugging without becoming part of the application's behavior. + +Log messages should answer: + +- What happened? +- Why did it happen? +- What should the developer investigate next? + +Avoid excessive or repetitive logging. + +--- + +# Documentation + +Document architecture rather than implementation details. + +Prefer: + +- ADRs +- specifications +- interfaces +- module responsibilities + +Avoid comments that merely repeat what the code already expresses. + +Comments should explain *why*, not *what*. + +--- + +# Naming + +Choose names that reflect business concepts rather than implementation details. + +Prefer: + +- ConversationTurn +- CaptureConnector +- ManifestGenerator + +Avoid: + +- TempData +- Helper +- Utils +- Stuff + +Names should communicate intent. + +--- + +# Interfaces + +Define interfaces at architectural boundaries. + +Examples include: + +- Capture Connector +- Processor +- Serializer + +Interfaces should remain stable even when implementations change. + +--- + +# Dependencies + +Minimize external dependencies. + +Before adding a dependency, ask: + +- Does the standard library already solve this? +- Does this dependency simplify architecture? +- Is it actively maintained? +- Can the dependency be removed later without significant effort? + +Favor long-lived, well-supported libraries. + +--- + +# File Formats + +Project Thoth prefers open formats. + +Preferred formats include: + +- Markdown +- YAML +- JSON +- Plain text + +Avoid introducing proprietary formats unless there is a compelling architectural reason. + +--- + +# Testing + +Testing should occur at multiple levels. + +Examples: + +- Unit tests +- Integration tests +- End-to-end workflow tests + +Where practical, preserve representative fixtures for regression testing. + +Regression tests are especially important for parsers and capture connectors. + +--- + +# AI-Assisted Development + +AI is a development assistant, not an architectural authority. + +AI-generated code should be reviewed for: + +- correctness +- simplicity +- architectural alignment +- maintainability + +AI should implement established architecture rather than invent new architecture during implementation. + +--- + +# Work Orders + +Implementation work should be described using Work Orders. + +A Work Order should include: + +- context +- objective +- constraints +- deliverables +- acceptance criteria +- architectural references + +Work Orders should be executable with minimal additional explanation. + +--- + +# Refactoring + +Refactoring should preserve observable behavior while improving internal structure. + +Typical reasons include: + +- simplifying design +- improving readability +- reducing duplication +- improving testability +- aligning implementation with updated architecture + +Refactoring should not introduce unrelated feature work. + +--- + +# Code Reviews + +Reviews should evaluate: + +1. Architectural alignment +2. Correctness +3. Readability +4. Maintainability +5. Simplicity + +The primary question is not: + +> "Does it work?" + +Instead ask: + +> "Will another developer understand this in two years?" + +--- + +# Performance + +Optimize for clarity first. + +Optimize for performance when: + +- measurement identifies a bottleneck +- scalability requires it +- architectural goals demand it + +Avoid premature optimization. + +--- + +# Security + +Treat external input as untrusted. + +Validate: + +- user input +- downloaded content +- file paths +- URLs +- serialized data + +Avoid embedding secrets in source code or repositories. + +--- + +# Backward Compatibility + +When changing canonical formats: + +- preserve compatibility where practical +- document breaking changes +- update specifications +- record architectural decisions in ADRs + +--- + +# Technical Debt + +Technical debt should be visible. + +Known debt should be documented rather than ignored. + +Temporary solutions should include: + +- rationale +- limitations +- expected future resolution + +--- + +# Definition of Done + +Work is considered complete when: + +- implementation satisfies the Work Order +- acceptance criteria are met +- architecture remains consistent +- documentation is updated if required +- tests pass (where applicable) +- no known regressions have been introduced + +Completion is measured by quality, not by the number of lines of code written. + +--- + +# Coding Philosophy + +Project Thoth values software that is: + +- understandable +- deterministic +- modular +- portable +- maintainable +- testable + +The objective is to create software that remains useful long after the original implementation has evolved. + +--- + +# Final Principle + +Good software preserves knowledge. + +Great software preserves understanding. + +Every design decision, module, and line of code should make the project easier for the next developer—or the next AI assistant—to understand. \ No newline at end of file diff --git a/.thoth/50-roadmap.md b/.thoth/50-roadmap.md new file mode 100644 index 0000000..7a14195 --- /dev/null +++ b/.thoth/50-roadmap.md @@ -0,0 +1,419 @@ +# Project Thoth Roadmap + +**Version:** 1.0 +**Status:** Living Document + +--- + +# Purpose + +This roadmap describes the long-term evolution of Project Thoth. + +Unlike `00-project-state.md`, which describes the project's current implementation, this document describes the intended direction of the platform. + +The roadmap is intentionally high-level. + +Implementation details belong in Work Orders, ADRs, and project-specific documentation. + +--- + +# Vision + +Project Thoth will become a platform for preserving, organizing, processing, and retrieving human knowledge independent of any AI provider or software vendor. + +Its primary objective is to create a durable knowledge architecture where information can be captured once and continuously reprocessed as technology evolves. + +--- + +# Guiding Principles + +The roadmap follows several principles: + +- Build vertically before horizontally. +- Preserve before processing. +- Prefer deterministic systems. +- Keep platform-specific code isolated. +- Favor reusable processors over application-specific logic. +- Introduce AI only after preservation is complete. + +--- + +# Phase 1 — Foundation + +## Objective + +Establish the architectural and technical foundations of Project Thoth. + +### Deliverables + +- Reference Architecture +- Project structure +- ADR framework +- Repository standards +- Canonical specifications +- `.thoth` project self-description +- Initial Work Order framework + +### Success Criteria + +Project architecture is stable and documented. + +--- + +# Phase 2 — Capture MVP + +## Objective + +Capture conversations from ChatGPT into canonical Project Thoth artifacts. + +### Deliverables + +- Browser extension +- ChatGPT connector +- Conversation discovery +- DOM-to-Markdown transformation +- Markdown serialization +- Download support +- Basic user feedback + +### Success Criteria + +Conversations can be captured with high fidelity into `conversation.md`. + +--- + +# Phase 3 — Capture Framework + +## Objective + +Generalize the ChatGPT implementation into a reusable connector framework. + +### Deliverables + +- Capture Connector interface +- Canonical Conversation Model +- Shared HTML-to-Markdown engine +- Shared Markdown serializer +- Connector testing framework + +### Target Connectors + +- ChatGPT +- Claude +- Gemini +- Microsoft Copilot +- Open WebUI + +### Success Criteria + +New connectors primarily require implementation of platform-specific discovery. + +--- + +# Phase 4 — Processor Framework + +## Objective + +Build reusable processors that operate on canonical source artifacts. + +### Initial Processors + +- Metadata Generator +- Manifest Generator +- Harvest Generator + +### Future Processors + +- Relationship extraction +- Citation generation +- Cross-reference generation +- Knowledge graph generation + +### Success Criteria + +Processors are independent of capture connectors. + +--- + +# Phase 5 — Archive Management + +## Objective + +Create a durable archive for canonical artifacts. + +### Features + +- Archive organization +- Collection management +- Version tracking +- Attachment management +- Integrity validation + +### Success Criteria + +The archive becomes the canonical source of knowledge. + +--- + +# Phase 6 — Knowledge Repository + +## Objective + +Organize captured knowledge into a coherent repository. + +### Features + +- Cross-linking +- Collections +- Relationships +- Metadata indexing +- Search preparation + +### Success Criteria + +Knowledge can be navigated independently of its original source platform. + +--- + +# Phase 7 — Search + +## Objective + +Provide deterministic search across the repository. + +### Features + +- Full-text search +- Metadata search +- Collection search +- Relationship search +- Citation search + +### Success Criteria + +Users can reliably locate preserved knowledge without AI assistance. + +--- + +# Phase 8 — AI Integration + +## Objective + +Enable AI systems to consume the repository. + +### Features + +- Local LLM integration +- Context packaging +- Prompt generation +- Retrieval interfaces +- Knowledge injection + +### Target Platforms + +- Ollama +- Open WebUI +- ChatGPT +- Claude +- Gemini + +### Success Criteria + +AI systems become consumers of the archive rather than owners of knowledge. + +--- + +# Phase 9 — Knowledge Operations + +## Objective + +Support organizational knowledge management. + +### Features + +- Team repositories +- Shared collections +- Governance +- Retention policies +- Workflow automation +- Knowledge health metrics + +### Success Criteria + +Organizations can manage institutional knowledge using Project Thoth. + +--- + +# Phase 10 — Project Thoth Platform + +## Objective + +Transform Project Thoth into a complete knowledge platform. + +### Features + +- Native applications +- Plugin architecture +- Processor marketplace +- Connector marketplace +- Automation framework +- Scheduled processing +- Background services + +### Success Criteria + +Project Thoth becomes an extensible ecosystem. + +--- + +# Parallel Work Streams + +Several work streams evolve throughout the roadmap. + +--- + +## Capture Connectors + +Examples include: + +- ChatGPT +- Claude +- Gemini +- Microsoft Copilot +- Open WebUI +- Cursor +- GitHub Copilot Chat +- Future AI platforms + +--- + +## Processors + +Examples include: + +- Metadata +- Manifest +- Harvest +- Summaries +- Search indexing +- Embeddings +- Relationship analysis + +--- + +## Specifications + +Examples include: + +- Conversation Specification +- Metadata Specification +- Manifest Specification +- Harvest Specification +- Package Specification + +--- + +## Applications + +Potential applications include: + +- Browser Capture +- Desktop Archive Manager +- Knowledge Explorer +- Search Portal +- Processor Runner +- Administration Console + +--- + +# Future Capabilities + +Potential future capabilities include: + +- Automatic conversation synchronization +- Email capture +- Document ingestion +- PDF processing +- Audio transcription +- Video indexing +- OCR +- Knowledge graph visualization +- Timeline reconstruction +- Citation management +- Research workspaces +- Multi-user collaboration + +These features are intentionally deferred until the core preservation architecture is mature. + +--- + +# Non-Goals + +Project Thoth is not intended to become: + +- another chatbot +- another note-taking application +- another vector database +- another proprietary knowledge silo + +These technologies may be integrated where appropriate, but they are not the project's primary purpose. + +--- + +# Milestone Philosophy + +Each milestone should deliver independently valuable functionality. + +The project should remain useful even if later phases are never completed. + +Earlier phases should minimize assumptions about later implementation choices. + +--- + +# Success Measures + +Progress should be measured by: + +- Preservation fidelity +- Architectural stability +- Deterministic processing +- Platform independence +- Ease of extension +- Long-term maintainability + +Feature count is not a primary measure of success. + +--- + +# Living Roadmap + +This roadmap is expected to evolve. + +New phases may be added. + +Existing phases may be reordered. + +However, changes should preserve the project's guiding philosophy: + +1. Capture. +2. Preserve. +3. Process. +4. Organize. +5. Retrieve. +6. Reason. + +Reasoning is the final stage—not the first. + +--- + +# Final Vision + +Project Thoth is not simply a collection of tools. + +It is an architecture for preserving human reasoning. + +The long-term goal is to ensure that knowledge remains accessible, portable, and understandable regardless of which AI systems, software platforms, or technologies come and go. + +Technology will continue to evolve. + +The archive should not have to. \ No newline at end of file diff --git a/README.md b/README.md new file mode 100644 index 0000000..088e6c3 --- /dev/null +++ b/README.md @@ -0,0 +1,7 @@ +They should be able to answer: + +What is Project Thoth? +Why does it exist? +Is this something I want to use? +How do I build it? +How do I contribute? \ No newline at end of file diff --git a/applications/chatgpt-capture/README.md b/applications/chatgpt-capture/README.md new file mode 100644 index 0000000..2c73939 --- /dev/null +++ b/applications/chatgpt-capture/README.md @@ -0,0 +1,33 @@ +# Project Thoth Capture + +This folder contains a minimal Manifest V3 browser extension skeleton for Chrome and Edge. + +## Local installation + +1. Open Chrome or Edge. +2. Navigate to `chrome://extensions` or `edge://extensions`. +3. Enable "Developer mode". +4. Click "Load unpacked". +5. Select the `applications/chatgpt-capture` folder in this repository. + +## Current behavior + +- The toolbar button is visible in the browser toolbar. +- Clicking it checks whether the active tab is a ChatGPT page. +- On a supported ChatGPT tab, the extractor runs in the page context and sends a capture payload back to the service worker. +- On unsupported tabs, the extension logs a simple warning and does not inject the script. +- Future extraction logic can be expanded in `src/chatgptExtractor.js` and `src/background.js`. + +## How to test against a live ChatGPT page + +1. Load the unpacked extension from this folder in Chrome or Edge developer mode. +2. Open a real ChatGPT conversation in a tab. +3. Open the extension's service worker console: + - Visit `chrome://extensions` or `edge://extensions`. + - Find "Project Thoth Capture" and click "service worker" or "inspect views". +4. Click the extension toolbar button while the ChatGPT tab is active. +5. Inspect the console for: + - `ChatGPT extractor payload:` + - `Received capture payload:` +6. If the page is detected as a conversation, the payload should include `title`, `url`, `capturedAt`, and `messages`. + - If no messages are found, the payload will include an `error` field instead. diff --git a/applications/chatgpt-capture/icons/icon128.png b/applications/chatgpt-capture/icons/icon128.png new file mode 100644 index 0000000..b3a4252 --- /dev/null +++ b/applications/chatgpt-capture/icons/icon128.png @@ -0,0 +1 @@ +placeholder \ No newline at end of file diff --git a/applications/chatgpt-capture/icons/icon16.png b/applications/chatgpt-capture/icons/icon16.png new file mode 100644 index 0000000..b3a4252 --- /dev/null +++ b/applications/chatgpt-capture/icons/icon16.png @@ -0,0 +1 @@ +placeholder \ No newline at end of file diff --git a/applications/chatgpt-capture/icons/icon48.png b/applications/chatgpt-capture/icons/icon48.png new file mode 100644 index 0000000..b3a4252 --- /dev/null +++ b/applications/chatgpt-capture/icons/icon48.png @@ -0,0 +1 @@ +placeholder \ No newline at end of file diff --git a/applications/chatgpt-capture/manifest.json b/applications/chatgpt-capture/manifest.json new file mode 100644 index 0000000..08adaa7 --- /dev/null +++ b/applications/chatgpt-capture/manifest.json @@ -0,0 +1,41 @@ +{ + "manifest_version": 3, + "name": "Project Thoth Capture", + "description": "Capture ChatGPT conversations for Project Thoth.", + "version": "0.1.0", + "permissions": ["activeTab", "scripting", "downloads"], + "host_permissions": [ + "https://chatgpt.com/*", + "https://chat.openai.com/*" + ], + "action": { + "default_title": "Save to Project Thoth", + "default_icon": { + "16": "icons/icon16.png", + "48": "icons/icon48.png", + "128": "icons/icon128.png" + } + }, + "background": { + "service_worker": "src/background.js", + "type": "module" + }, + "web_accessible_resources": [ + { + "resources": [ + "src/chatgptConnectorRuntime.js", + "src/connectors/chatgpt/chatgptConnector.js", + "src/connectors/chatgpt/chatgptDiscovery.js", + "src/connectors/chatgpt/chatgptPreprocessor.js", + "src/connectors/chatgpt/chatgptTurnRenderer.js", + "src/core/conversationModel.js", + "src/core/htmlToMarkdown.js", + "src/core/markdownSerializer.js" + ], + "matches": [ + "https://chatgpt.com/*", + "https://chat.openai.com/*" + ] + } + ] +} diff --git a/applications/chatgpt-capture/src/background.js b/applications/chatgpt-capture/src/background.js new file mode 100644 index 0000000..bee2683 --- /dev/null +++ b/applications/chatgpt-capture/src/background.js @@ -0,0 +1,162 @@ +import { createConversationFilename } from "./filename.js"; +import { serializeConversationToMarkdown } from "./core/markdownSerializer.js"; + +const CHATGPT_HOSTS = ["https://chatgpt.com/", "https://chat.openai.com/"]; +const pendingDownloadIds = new Set(); + +const USER_MESSAGES = { + NOT_ON_CHATGPT: "Open a ChatGPT conversation before saving to Project Thoth.", + NO_CONVERSATION: "No ChatGPT conversation was detected on this page.", + CAPTURE_FAILED: "Capture failed. Try scrolling through the conversation and capturing again.", + DOWNLOAD_FAILED: "The conversation was captured, but the file could not be saved.", + SUCCESS: "Saved to Downloads." +}; + +function showStatus(tabId, message, badgeText = "") { + if (typeof tabId === "number") { + chrome.action.setTitle({ tabId, title: message }); + chrome.action.setBadgeText({ tabId, text: badgeText }); + + if (badgeText) { + setTimeout(() => { + chrome.action.setBadgeText({ tabId, text: "" }); + chrome.action.setTitle({ tabId, title: "Save to Project Thoth" }); + }, 3000); + } + } else { + console.log(`Status: ${message}`); + } +} + +function isChatGPTUrl(url) { + if (!url) { + return false; + } + + return CHATGPT_HOSTS.some((host) => url.startsWith(host)); +} + +chrome.action.onClicked.addListener(async (tab) => { + const tabUrl = tab?.url || "unknown"; + const tabTitle = tab?.title || "unknown"; + + console.log("Project Thoth capture clicked"); + console.log(`Tab URL: ${tabUrl}`); + console.log(`Tab title: ${tabTitle}`); + + if (!tab?.id || !isChatGPTUrl(tabUrl)) { + console.warn("Capture skipped: active tab is not a ChatGPT page."); + showStatus(tab.id, USER_MESSAGES.NOT_ON_CHATGPT, "ERR"); + return; + } + + try { + console.log("[Thoth] capture requested"); + await chrome.scripting.executeScript({ + target: { tabId: tab.id }, + func: async (runtimeUrl) => { + const { captureChatGPTConversation } = await import(runtimeUrl); + await captureChatGPTConversation(); + }, + args: [chrome.runtime.getURL("src/chatgptConnectorRuntime.js")] + }); + + console.log("Injected ChatGPT capture script."); + } catch (error) { + console.error("Failed to inject capture script:", error); + showStatus(tab.id, USER_MESSAGES.CAPTURE_FAILED, "ERR"); + } +}); + +chrome.runtime.onMessage.addListener((message, sender) => { + const tabId = sender?.tab?.id; + + if (message?.type === "CHATGPT_CAPTURE_STATUS") { + const messageText = message?.payload?.message || "Preparing conversation for capture..."; + showStatus(tabId, messageText, "..."); + return; + } + + if (message?.type === "CHATGPT_CAPTURE_PAYLOAD") { + console.log("Received capture payload:", message.payload); + console.log("Sender tab URL:", sender?.tab?.url || "unknown"); + + if (message.payload?.error || !Array.isArray(message.payload?.messages) || message.payload.messages.length === 0) { + console.warn("No ChatGPT conversation was detected on this page.", message.payload); + showStatus(tabId, USER_MESSAGES.NO_CONVERSATION, "ERR"); + return; + } + + console.log("[Thoth] extraction completed"); + + try { + const markdown = serializeConversationToMarkdown({ + title: message.payload.title, + sourcePlatform: message.payload.sourcePlatform, + url: message.payload.url, + capturedAt: message.payload.capturedAt, + warning: message.payload.warning, + turns: (message.payload.messages || []).map((messageItem) => ({ + role: messageItem.role, + markdown: messageItem.content, + captureStatus: messageItem.captureStatus || "captured" + })) + }); + const filename = createConversationFilename(message.payload); + const blob = new Blob([markdown], { type: "text/markdown;charset=utf-8" }); + const downloadUrl = + typeof URL.createObjectURL === "function" + ? URL.createObjectURL(blob) + : `data:text/markdown;charset=utf-8,${encodeURIComponent(markdown)}`; + + console.log("Normalized markdown output:\n" + markdown); + console.log("Generated filename:", filename); + console.log("[Thoth] markdown generated"); + console.log("[Thoth] download requested"); + + chrome.downloads.download( + { + url: downloadUrl, + filename, + saveAs: false, + conflictAction: "uniquify" + }, + (downloadId) => { + if (typeof URL.revokeObjectURL === "function" && downloadUrl.startsWith("blob:")) { + URL.revokeObjectURL(downloadUrl); + } + + if (chrome.runtime.lastError || !downloadId) { + console.error( + "[Thoth] download failed:", + chrome.runtime.lastError?.message || "No download ID was returned." + ); + showStatus(tabId, USER_MESSAGES.DOWNLOAD_FAILED, "ERR"); + return; + } + + console.log("Markdown download started with ID:", downloadId); + pendingDownloadIds.add(downloadId); + showStatus(tabId, USER_MESSAGES.SUCCESS, "OK"); + } + ); + } catch (error) { + console.error("Failed to normalize capture payload:", error); + showStatus(tabId, USER_MESSAGES.CAPTURE_FAILED, "ERR"); + } + } +}); + +chrome.downloads.onChanged.addListener((delta) => { + if (!pendingDownloadIds.has(delta.id)) { + return; + } + + if (delta.state?.current === "complete") { + console.log(`[Thoth] download completed: ${delta.id}`); + pendingDownloadIds.delete(delta.id); + } else if (delta.error?.current) { + console.error(`[Thoth] download failed: ${delta.error.current}`); + pendingDownloadIds.delete(delta.id); + } +}); diff --git a/applications/chatgpt-capture/src/chatgptConnectorRuntime.js b/applications/chatgpt-capture/src/chatgptConnectorRuntime.js new file mode 100644 index 0000000..922ff7c --- /dev/null +++ b/applications/chatgpt-capture/src/chatgptConnectorRuntime.js @@ -0,0 +1,49 @@ +import { captureConversation } from "./connectors/chatgpt/chatgptConnector.js"; + +export async function captureChatGPTConversation() { + console.log("[Thoth] connector runtime loaded"); + + try { + const result = await captureConversation({ + title: document.title, + url: window.location.href, + capturedAt: new Date().toISOString() + }); + + if (typeof chrome !== "undefined" && chrome.runtime && chrome.runtime.sendMessage) { + chrome.runtime.sendMessage({ + type: "CHATGPT_CAPTURE_PAYLOAD", + payload: { + sourcePlatform: result.conversation.sourcePlatform, + title: result.conversation.title, + url: result.conversation.url, + capturedAt: result.conversation.capturedAt, + warning: result.conversation.warning, + messages: result.conversation.turns.map((turn) => ({ + role: turn.role, + content: turn.markdown, + captureStatus: turn.captureStatus, + turnIndex: turn.turnIndex + })), + debug: result.conversation.debug + } + }); + } + } catch (error) { + console.error("[Thoth] extraction failed:", error); + + if (typeof chrome !== "undefined" && chrome.runtime && chrome.runtime.sendMessage) { + chrome.runtime.sendMessage({ + type: "CHATGPT_CAPTURE_PAYLOAD", + payload: { + sourcePlatform: "ChatGPT", + title: document.title || "ChatGPT Conversation", + url: window.location.href, + capturedAt: new Date().toISOString(), + messages: [], + error: error.message || "ChatGPT capture failed." + } + }); + } + } +} diff --git a/applications/chatgpt-capture/src/chatgptExtractor.js b/applications/chatgpt-capture/src/chatgptExtractor.js new file mode 100644 index 0000000..48bdc11 --- /dev/null +++ b/applications/chatgpt-capture/src/chatgptExtractor.js @@ -0,0 +1,994 @@ +(function () { + const MESSAGE_CONTAINER_SELECTORS = [ + '[data-message-author-role]', + '[data-testid*="message"]', + '[data-testid*="conversation-turn"]', + '[data-testid*="assistant"]', + '[data-testid*="user"]', + '[data-message-id]', + '.message', + '.conversation-turn', + '.assistant-message', + '.user-message' + ]; + + const SKIP_SELECTORS = [ + 'nav', + 'aside', + 'header', + 'footer', + 'form', + 'input', + 'textarea', + 'button', + '[role="button"]', + '[role="textbox"]', + '[contenteditable="true"]', + 'select', + 'svg', + 'script', + 'style' + ]; + + const UI_CHROME_SELECTORS = [ + 'button', + '[role="button"]', + '[aria-label*="copy" i]', + '[aria-label*="clipboard" i]', + '[aria-label*="feedback" i]', + '[data-testid*="copy" i]', + '[data-testid*="feedback" i]', + '[data-testid*="toolbar" i]', + '[data-testid*="model" i]', + '[data-testid*="timestamp" i]', + '[class*="copy" i]', + '[class*="feedback" i]', + '[class*="toolbar" i]', + '[class*="timestamp" i]' + ]; + + function normalizeWhitespace(text) { + if (!text) { + return ""; + } + + return text + .replace(/\u00a0/g, " ") + .replace(/[ \t]+\n/g, "\n") + .replace(/\n{3,}/g, "\n\n") + .replace(/[ \t]{2,}/g, " ") + .trim(); + } + + function isHiddenOrInvisible(node) { + if (!node || typeof node.matches !== "function") { + return true; + } + + if (node.matches('[aria-hidden="true"], [hidden], .hidden, .sr-only')) { + return true; + } + + if (typeof node.getBoundingClientRect !== "function") { + return false; + } + + try { + const rect = node.getBoundingClientRect(); + return rect.width === 0 && rect.height === 0; + } catch (error) { + return false; + } + } + + function hasMeaningfulText(node) { + if (!node) { + return false; + } + + const text = (node.textContent || "").replace(/\s+/g, "").trim(); + return text.length >= 6; + } + + function isChromeUi(node) { + if (!node || typeof node.matches !== "function") { + return true; + } + + if (node.matches(SKIP_SELECTORS.join(", ")) || node.closest(SKIP_SELECTORS.join(", "))) { + return true; + } + + return false; + } + + function getConversationTitle() { + const titleSelectors = [ + '[data-testid*="title"]', + '[data-testid*="conversation-title"]', + 'main h1', + 'main h2', + '[role="heading"]' + ]; + + for (const selector of titleSelectors) { + const element = document.querySelector(selector); + if (element && hasMeaningfulText(element)) { + return normalizeWhitespace(element.textContent || ""); + } + } + + return document.title || "ChatGPT Conversation"; + } + + function getConversationRoot() { + const possibleRoots = [ + document.querySelector('main[role="main"]'), + document.querySelector('main'), + document.querySelector('[role="main"]'), + document.querySelector('[data-testid*="conversation"]'), + document.querySelector('[data-testid*="messages"]'), + document.querySelector('[data-testid*="thread"]'), + document.querySelector('[data-testid*="chat"]'), + document.querySelector('[class*="conversation"]'), + document.querySelector('[class*="messages"]'), + document.querySelector('[class*="thread"]'), + document.body + ]; + + return possibleRoots.find(Boolean) || document.body; + } + + function getElementFingerprint(node) { + return normalizeWhitespace((node.textContent || "").replace(/\s+/g, " ")).toLowerCase(); + } + + function isShellContainer(node, root) { + if (!node || !root || node === root) { + return false; + } + + const label = [ + node.getAttribute("aria-label"), + node.getAttribute("data-testid"), + node.className, + node.textContent || "" + ].join(" ").toLowerCase(); + + const shellHints = [ + "sidebar", + "history", + "nav", + "navigation", + "composer", + "input", + "textarea", + "search", + "account", + "project", + "share", + "settings", + "footer", + "menu", + "toolbar", + "new chat", + "skip to content", + "window.__oai_loghtml" + ]; + + if (shellHints.some((hint) => label.includes(hint))) { + return true; + } + + const shellSelector = [ + 'nav', + 'aside', + 'header', + 'footer', + 'form', + 'input', + 'textarea', + 'button', + '[role="button"]', + '[role="textbox"]', + '[role="navigation"]', + '[role="menu"]', + '[role="tooltip"]', + '[contenteditable="true"]', + '[data-testid*="sidebar"]', + '[data-testid*="history"]', + '[data-testid*="composer"]', + '[data-testid*="input"]', + '[data-testid*="toolbar"]', + '[data-testid*="share"]', + '[data-testid*="account"]' + ].join(", "); + + if (node.matches(shellSelector)) { + return true; + } + + const shellAncestor = node.closest(shellSelector); + return Boolean(shellAncestor && shellAncestor !== root && root.contains(shellAncestor)); + } + + function looksLikeTranscriptNode(node, root) { + if (!node || !root || node === root) { + return false; + } + + if (isChromeUi(node) || isHiddenOrInvisible(node) || isShellContainer(node, root)) { + return false; + } + + const text = normalizeWhitespace(node.textContent || ""); + if (!text) { + return false; + } + + if (/window\.__oai_loghtml|/i.test(text)) { + return false; + } + + const label = [ + node.getAttribute("data-testid"), + node.getAttribute("aria-label"), + node.className, + text + ].join(" ").toLowerCase(); + + const shellWords = /sidebar|history|composer|input|toolbar|settings|share|footer|menu|account|project|navigation|new chat|skip to content/i.test(label); + if (shellWords) { + return false; + } + + const roleAncestor = node.closest('[data-message-author-role], [data-testid*="message"], [data-testid*="conversation"], [data-testid*="assistant"], [data-testid*="user"], [data-testid*="turn"]'); + const hasAncestorRoleSignal = Boolean(roleAncestor && root.contains(roleAncestor) && roleAncestor !== node); + const messageSignals = /message|conversation|turn|assistant|user|chatgpt|you\s+said|chatgpt\s+said/i.test(label); + const tagName = (node.tagName || "").toLowerCase(); + const isBlockLike = ["div", "section", "article", "p", "li", "td", "th", "pre", "blockquote"].includes(tagName); + const childCount = Array.from(node.children || []).filter((child) => child.nodeType === Node.ELEMENT_NODE).length; + const hasBlockStructure = childCount <= 12 && (node.querySelector("p, li, pre, blockquote, table") || text.length >= 8); + const hasMeaningfulLength = text.length >= 8 && text.length <= 4000; + + return (messageSignals || hasAncestorRoleSignal || (isBlockLike && hasBlockStructure && hasMeaningfulLength)) && !shellWords; + } + + function splitIntoTranscriptTurns(text, fallbackRole) { + const normalized = normalizeWhitespace(text || ""); + if (!normalized) { + return []; + } + + const parts = normalized.split(/(?=\b(?:you|user|assistant|chatgpt)\b)/i).map((part) => part.trim()).filter(Boolean); + if (parts.length <= 1) { + return [{ role: fallbackRole || "unknown", content: normalized }]; + } + + return parts.map((part) => { + const roleText = part.match(/^(you|user|assistant|chatgpt)\b/i); + const role = roleText ? (roleText[1].toLowerCase().includes("assistant") || roleText[1].toLowerCase() === "chatgpt" ? "assistant" : roleText[1].toLowerCase() === "user" || roleText[1].toLowerCase() === "you" ? "user" : "unknown") : (fallbackRole || "unknown"); + const content = part.replace(/^(you|user|assistant|chatgpt)\b(?:\s*(?:said|:|-))?\s*/i, "").trim(); + return { + role, + content: content || part + }; + }).filter((item) => Boolean(item.content)); + } + + function collectMessageCandidates() { + const root = getConversationRoot(); + const candidates = []; + const shellSkipped = []; + const duplicateSkipped = []; + const orphanChildBlocksSkipped = []; + + const possibleNodes = Array.from(root.querySelectorAll("*")); + possibleNodes.forEach((element) => { + if (!element || !hasMeaningfulText(element)) { + return; + } + + if (isChromeUi(element) || isHiddenOrInvisible(element)) { + return; + } + + if (isShellContainer(element, root)) { + shellSkipped.push(element); + return; + } + + if (!looksLikeTranscriptNode(element, root)) { + return; + } + + const text = getElementFingerprint(element); + if (!text) { + return; + } + + candidates.push(element); + }); + + const orderedCandidates = candidates.sort((left, right) => { + const leftDepth = left && left.tagName !== "BODY" ? (left.parentElement ? Array.from(left.parentElement.children).indexOf(left) : 0) : 0; + const rightDepth = right && right.tagName !== "BODY" ? (right.parentElement ? Array.from(right.parentElement.children).indexOf(right) : 0) : 0; + if (leftDepth !== rightDepth) { + return leftDepth - rightDepth; + } + + const leftTextLength = (left.textContent || "").length; + const rightTextLength = (right.textContent || "").length; + return rightTextLength - leftTextLength; + }); + + const deduped = []; + orderedCandidates.forEach((candidate) => { + const hasAncestorCandidate = deduped.some((existing) => existing !== candidate && existing.contains(candidate)); + if (hasAncestorCandidate) { + orphanChildBlocksSkipped.push(candidate); + return; + } + + const fingerprint = getElementFingerprint(candidate); + const isDuplicate = deduped.some((existing) => { + const existingFingerprint = getElementFingerprint(existing); + const sameText = existingFingerprint === fingerprint; + const nested = existing.contains(candidate) || candidate.contains(existing); + return sameText && nested; + }); + + if (isDuplicate) { + duplicateSkipped.push(candidate); + return; + } + + deduped.push(candidate); + }); + + return { + candidates: deduped, + debug: { + candidateContainersFound: orderedCandidates.length, + afterShellFilter: deduped.length, + afterDeduplication: deduped.length, + skippedShellContainers: shellSkipped.length, + skippedDuplicateContainers: duplicateSkipped.length, + orphanChildBlocksSkipped: orphanChildBlocksSkipped.length + } + }; + } + + function detectMessageRole(element) { + const parts = []; + let current = element; + + while (current && current !== document.body) { + parts.push([ + current.getAttribute("data-message-author-role"), + current.getAttribute("data-testid"), + current.getAttribute("aria-label"), + current.className + ].join(" ")); + current = current.parentElement; + } + + const roleText = parts.join(" ").toLowerCase(); + + if (roleText.includes("assistant") || roleText.includes("chatgpt")) { + return "assistant"; + } + + if (roleText.includes("user") || roleText.includes("you")) { + return "user"; + } + + return "unknown"; + } + + function classifyContainer(container) { + const role = detectMessageRole(container); + const text = normalizeWhitespace(container.textContent || ""); + const hasStrongSignals = /assistant|chatgpt|you|user/i.test(container.getAttribute("data-testid") || "") || /assistant|chatgpt|you|user/i.test(container.className || ""); + + return { + role, + type: hasStrongSignals ? "conversation" : "unknown", + rawText: text, + htmlSnippet: (container.outerHTML || "").slice(0, 240) + }; + } + + function extractContainerFallback(container) { + const text = normalizeWhitespace(container.textContent || ""); + return { + role: detectMessageRole(container), + rawText: text, + htmlSnippet: (container.outerHTML || "").slice(0, 240) + }; + } + + function findConversationContainers() { + const discovery = collectMessageCandidates(); + const containers = discovery.candidates.map((container) => ({ + element: container, + ...classifyContainer(container) + })); + + return { + containers, + debug: { + ...discovery.debug, + candidateContainersFound: discovery.debug.candidateContainersFound, + afterShellFilter: discovery.debug.afterShellFilter, + afterDeduplication: discovery.debug.afterDeduplication, + skippedShellContainers: discovery.debug.skippedShellContainers, + skippedDuplicateContainers: discovery.debug.skippedDuplicateContainers, + orphanChildBlocksSkipped: discovery.debug.orphanChildBlocksSkipped + } + }; + } + + function getCodeLanguage(node) { + const className = (node.className || "").toString(); + const languageMatch = className.match(/language-([a-z0-9_+-]+)/i) || className.match(/lang-([a-z0-9_+-]+)/i); + const dataLanguage = node.getAttribute("data-language") || node.getAttribute("data-lang"); + + if (languageMatch) { + return languageMatch[1]; + } + + return dataLanguage || ""; + } + + function normalizeInlineText(text) { + return (text || "") + .replace(/\u00a0/g, " ") + .replace(/\s+/g, " ") + .replace(/\s+([,.;:!?])/g, "$1") + .trim(); + } + + function escapeMarkdownLinkText(text) { + return String(text || "") + .replace(/\[/g, "\\[") + .replace(/\]/g, "\\]") + .trim(); + } + + function escapeMarkdownLinkUrl(url) { + return String(url || "") + .replace(/\)/g, "%29") + .trim(); + } + + function isSafeHref(href) { + if (!href) { + return false; + } + + const safeHref = href.trim(); + if (!safeHref || /^javascript:/i.test(safeHref) || /^data:/i.test(safeHref) || /^blob:/i.test(safeHref)) { + return false; + } + + return /^https?:\/\//i.test(safeHref) || /^mailto:/i.test(safeHref); + } + + function linkToMarkdown(anchorNode) { + const href = anchorNode.getAttribute("href") || ""; + const textContent = normalizeInlineText(childrenToInlineMarkdown(anchorNode)); + const visibleText = textContent || normalizeInlineText(anchorNode.textContent || ""); + const label = visibleText || href; + + if (!isSafeHref(href)) { + return visibleText; + } + + return `[${escapeMarkdownLinkText(label)}](${escapeMarkdownLinkUrl(href)})`; + } + + function isInlineElement(node) { + return Boolean( + node && + node.nodeType === Node.ELEMENT_NODE && + node.matches("strong, b, em, i, code, span, a, small, sub, sup, u, s") + ); + } + + function isBlockElement(node) { + return Boolean( + node && + node.nodeType === Node.ELEMENT_NODE && + node.matches("p, ul, ol, li, pre, blockquote, h1, h2, h3, h4, h5, h6, table, tr, td, th, div, section, article, main, hr") + ); + } + + function escapeTableCell(value) { + return String(value || "") + .replace(/\|/g, "\\|") + .replace(/\r?\n/g, "
") + .replace(/\s+/g, " ") + .trim(); + } + + function tableToMarkdown(tableNode) { + const rows = Array.from(tableNode.querySelectorAll("tr")); + if (rows.length === 0) { + return ""; + } + + const headerRow = rows.find((row) => row.querySelector("th")) || rows[0]; + const bodyRows = rows.filter((row) => row !== headerRow); + const headerCells = Array.from(headerRow.querySelectorAll("th, td")) + .map((cell) => escapeTableCell(childrenToInlineMarkdown(cell))) + .filter((cell) => cell.length > 0); + + if (headerCells.length === 0) { + return ""; + } + + const body = bodyRows.map((row) => { + const cells = Array.from(row.querySelectorAll("td")) + .map((cell) => escapeTableCell(childrenToInlineMarkdown(cell))) + .filter((cell) => cell.length > 0); + + if (cells.length === 0) { + return ""; + } + + return `| ${cells.join(" | ")} |`; + }).filter(Boolean); + + const headerLine = `| ${headerCells.join(" | ")} |`; + const separatorLine = `| ${headerCells.map(() => "---").join(" | ")} |`; + return [headerLine, separatorLine, ...body].join("\n"); + } + + function codeBlockToMarkdown(preNode) { + const codeNode = preNode.querySelector("code") || preNode; + const language = getCodeLanguage(codeNode); + const content = (codeNode.textContent || "") + .replace(/\r\n/g, "\n") + .replace(/\n{3,}/g, "\n\n") + .trim(); + + return [language ? `\`\`\`${language}` : "```", content, "```"].filter(Boolean).join("\n"); + } + + function childrenToInlineMarkdown(node) { + const parts = []; + + Array.from(node.childNodes).forEach((child) => { + if (child.nodeType === Node.TEXT_NODE) { + const text = normalizeInlineText(child.textContent || ""); + if (text) { + parts.push(text); + } + return; + } + + if (child.nodeType !== Node.ELEMENT_NODE) { + return; + } + + if (child.matches("strong, b")) { + parts.push(`**${childrenToInlineMarkdown(child)}**`); + } else if (child.matches("em, i")) { + parts.push(`*${childrenToInlineMarkdown(child)}*`); + } else if (child.matches("code")) { + parts.push(`\`${normalizeInlineText(child.textContent || "")}\``); + } else if (child.matches("a")) { + parts.push(linkToMarkdown(child)); + } else { + parts.push(nodeToMarkdown(child)); + } + }); + + return parts.join(" ").trim(); + } + + function listItemToMarkdown(liNode, depth = 0) { + const prefix = depth > 0 ? `${" ".repeat(depth)}` : ""; + const inlineContent = childrenToInlineMarkdown(liNode).trim(); + const nestedItems = Array.from(liNode.children || []).filter((child) => child.matches("ul, ol")); + + let content = inlineContent || normalizeInlineText(liNode.textContent || ""); + if (!content) { + return ""; + } + + const nestedMarkdown = nestedItems + .map((nestedList) => listToMarkdown(nestedList, depth + 1)) + .filter(Boolean) + .join("\n"); + + return [ + `${prefix}- ${content}`, + nestedMarkdown ? `${prefix} ${nestedMarkdown.replace(/\n/g, `\n${prefix} `)}` : "" + ].filter(Boolean).join("\n"); + } + + function listToMarkdown(listNode, depth = 0) { + const listType = listNode.tagName.toLowerCase(); + const items = Array.from(listNode.children || []).filter((child) => child.matches("li")); + + return items + .map((item, index) => { + const marker = listType === "ol" ? `${index + 1}. ` : "- "; + const inlineContent = childrenToInlineMarkdown(item).trim(); + const nestedItems = Array.from(item.children || []).filter((child) => child.matches("ul, ol")); + const nestedMarkdown = nestedItems + .map((nestedList) => listToMarkdown(nestedList, depth + 1)) + .filter(Boolean) + .join("\n"); + + const content = inlineContent || normalizeInlineText(item.textContent || ""); + const prefix = depth > 0 ? `${" ".repeat(depth)}` : ""; + + return [ + `${prefix}${marker}${content}`, + nestedMarkdown ? `${prefix}${" ".repeat(depth + 1)}${nestedMarkdown.replace(/\n/g, `\n${prefix}${" ".repeat(depth + 1)}`)}` : "" + ].filter(Boolean).join("\n"); + }) + .filter(Boolean) + .join("\n"); + } + + function blockquoteToMarkdown(blockquoteNode) { + return Array.from(blockquoteNode.childNodes) + .map((child) => { + if (child.nodeType === Node.TEXT_NODE) { + return `> ${normalizeInlineText(child.textContent || "")}`; + } + + if (child.nodeType !== Node.ELEMENT_NODE) { + return ""; + } + + const markdown = nodeToMarkdown(child); + return markdown + .split(/\n/) + .map((line) => (line ? `> ${line}` : ">")) + .join("\n"); + }) + .filter(Boolean) + .join("\n"); + } + + function cleanupMarkdownOutsideCodeFences(markdown) { + const lines = (markdown || "").split(/\r?\n/); + const cleaned = []; + let inCodeFence = false; + + lines.forEach((line) => { + const trimmed = line.trim(); + const isFence = /^```/.test(trimmed); + + if (isFence) { + inCodeFence = !inCodeFence; + cleaned.push(line); + return; + } + + if (inCodeFence) { + cleaned.push(line); + return; + } + + const normalizedLine = line.replace(/\s+([,.;:!?])/g, "$1"); + const punctuationOnly = /^[\s\p{P}\u2014\u2013]+$/u.test(normalizedLine); + + if (normalizedLine.trim() === "" || punctuationOnly) { + if (cleaned.length === 0 || cleaned[cleaned.length - 1].trim() !== "") { + cleaned.push(""); + } + return; + } + + cleaned.push(normalizedLine); + }); + + return cleaned.join("\n").replace(/\n{3,}/g, "\n\n").trim(); + } + + function blockChildrenToMarkdown(node) { + const parts = []; + let paragraphBuffer = []; + + const flushParagraph = () => { + const paragraph = paragraphBuffer.join(" ").trim(); + if (paragraph) { + parts.push(paragraph); + } + paragraphBuffer = []; + }; + + Array.from(node.childNodes).forEach((child) => { + if (child.nodeType === Node.TEXT_NODE) { + const text = normalizeInlineText(child.textContent || ""); + if (text) { + paragraphBuffer.push(text); + } + return; + } + + if (child.nodeType !== Node.ELEMENT_NODE) { + return; + } + + if (child.matches("p")) { + flushParagraph(); + parts.push(nodeToMarkdown(child)); + return; + } + + if (child.matches("ul, ol")) { + flushParagraph(); + parts.push(nodeToMarkdown(child)); + return; + } + + if (child.matches("pre")) { + flushParagraph(); + parts.push(nodeToMarkdown(child)); + return; + } + + if (child.matches("blockquote")) { + flushParagraph(); + parts.push(nodeToMarkdown(child)); + return; + } + + if (child.matches("table")) { + flushParagraph(); + parts.push(nodeToMarkdown(child)); + return; + } + + if (child.matches("h1, h2, h3, h4, h5, h6")) { + flushParagraph(); + parts.push(nodeToMarkdown(child)); + return; + } + + if (child.matches("br")) { + paragraphBuffer.push("\n"); + return; + } + + if (isInlineElement(child)) { + paragraphBuffer.push(childrenToInlineMarkdown(child)); + return; + } + + const nestedMarkdown = nodeToMarkdown(child); + if (nestedMarkdown) { + flushParagraph(); + parts.push(nestedMarkdown); + } + }); + + flushParagraph(); + + return parts.join("\n\n"); + } + + function nodeToMarkdown(node) { + if (!node) { + return ""; + } + + if (node.nodeType === Node.TEXT_NODE) { + return normalizeInlineText(node.textContent || ""); + } + + if (node.nodeType !== Node.ELEMENT_NODE) { + return ""; + } + + if (node.matches('script, style, svg, img, video, audio')) { + return ""; + } + + if (node.matches("pre")) { + return codeBlockToMarkdown(node); + } + + if (node.matches("code")) { + return `\`${normalizeInlineText(node.textContent || "")}\``; + } + + if (node.matches("br")) { + return "\n"; + } + + if (node.matches("blockquote")) { + return blockquoteToMarkdown(node); + } + + if (node.matches("table")) { + return tableToMarkdown(node); + } + + if (node.matches("ul, ol")) { + return listToMarkdown(node); + } + + if (node.matches("li")) { + return listItemToMarkdown(node); + } + + if (node.matches("h1, h2, h3, h4, h5, h6")) { + const level = parseInt(node.tagName.charAt(1), 10) || 1; + return `${"#".repeat(level)} ${childrenToInlineMarkdown(node)}`; + } + + if (node.matches("p")) { + return childrenToInlineMarkdown(node); + } + + if (isInlineElement(node)) { + return childrenToInlineMarkdown(node); + } + + return blockChildrenToMarkdown(node); + } + + function inlineChildrenToMarkdown(node) { + return childrenToInlineMarkdown(node); + } + + function containerToMarkdown(container) { + const element = container?.element || container; + if (!element) { + return ""; + } + + const rawText = normalizeWhitespace(element.textContent || ""); + const htmlText = (element.outerHTML || "").toLowerCase(); + const unsupportedKeywords = /(image generation|generated image|image|artifact|file upload|uploaded file|tool output|browser|code interpreter|analysis)/i; + + if (unsupportedKeywords.test(rawText) || unsupportedKeywords.test(htmlText)) { + const typeMatch = rawText.match(/image generation|artifact|file upload|tool output|browser|code interpreter|analysis/i); + const type = typeMatch ? typeMatch[0] : "unsupported content"; + const placeholder = `[Unsupported ChatGPT content: ${type}]`; + return rawText ? `${placeholder}\n\n${rawText}` : placeholder; + } + + const clone = prepareExtractableNode(element); + const markdown = extractNodeAsMarkdown(clone); + return markdown || rawText; + } + + function extractNodeAsMarkdown(node) { + if (!node) { + return ""; + } + + const markdown = nodeToMarkdown(node); + return cleanupMarkdownOutsideCodeFences(markdown); + } + + function prepareExtractableNode(node) { + const clone = node.cloneNode(true); + clone.querySelectorAll(UI_CHROME_SELECTORS.join(", ")).forEach((child) => child.remove()); + clone.querySelectorAll('script, style, svg, img, video, audio, [aria-hidden="true"], [hidden], .hidden, .sr-only').forEach((child) => child.remove()); + return clone; + } + + function extractMessageContent(element) { + return containerToMarkdown(element); + } + + function findMessageElements() { + return collectMessageCandidates().filter((element) => { + const content = extractMessageContent(element); + return Boolean(content); + }); + } + + function extractChatGPTConversation() { + const title = getConversationTitle(); + const discovery = findConversationContainers(); + const discoveredContainers = discovery.containers; + const parsedMessages = []; + const seenFingerprints = new Set(); + + discoveredContainers.forEach((container) => { + const rawContent = containerToMarkdown(container); + const content = normalizeWhitespace(rawContent || container.rawText || ""); + if (!content) { + return; + } + + const role = container.role || detectMessageRole(container.element); + const fingerprint = `${role}:${content.toLowerCase()}`; + if (seenFingerprints.has(fingerprint)) { + return; + } + + seenFingerprints.add(fingerprint); + parsedMessages.push({ + role, + content, + rawText: container.rawText, + type: container.type, + htmlSnippet: container.htmlSnippet + }); + }); + + const messages = parsedMessages + .map((message) => ({ + role: message.role, + content: message.content || message.rawText + })) + .filter((message) => Boolean(message.content)); + + const debug = { + turnContainersFound: discoveredContainers.length, + messagesReturned: messages.length, + duplicateTurnsSkipped: Math.max(0, discoveredContainers.length - parsedMessages.length), + orphanChildBlocksSkipped: discovery.debug.orphanChildBlocksSkipped || 0, + unknownTurnsReturned: parsedMessages.filter((message) => message.role === "unknown").length, + candidateContainersFound: discovery.debug.candidateContainersFound, + afterShellFilter: discovery.debug.afterShellFilter, + afterDeduplication: discovery.debug.afterDeduplication, + skippedShellContainers: discovery.debug.skippedShellContainers, + skippedDuplicateContainers: discovery.debug.skippedDuplicateContainers + }; + + if (messages.length === 0) { + return { + sourcePlatform: "ChatGPT", + title, + url: window.location.href, + capturedAt: new Date().toISOString(), + messages: [], + error: "No ChatGPT conversation messages were detected.", + debug + }; + } + + return { + sourcePlatform: "ChatGPT", + title, + url: window.location.href, + capturedAt: new Date().toISOString(), + messages, + debug + }; + } + + window.extractChatGPTConversation = extractChatGPTConversation; + + (async function () { + try { + const payload = await extractChatGPTConversation(); + console.log("ChatGPT extractor payload:", payload); + if (typeof chrome !== "undefined" && chrome.runtime && chrome.runtime.sendMessage) { + chrome.runtime.sendMessage({ + type: "CHATGPT_CAPTURE_PAYLOAD", + payload + }); + } + } catch (error) { + const fallbackPayload = { + sourcePlatform: "ChatGPT", + title: getConversationTitle(), + url: window.location.href, + capturedAt: new Date().toISOString(), + messages: [], + error: error.message || "ChatGPT extraction failed." + }; + + if (typeof chrome !== "undefined" && chrome.runtime && chrome.runtime.sendMessage) { + chrome.runtime.sendMessage({ + type: "CHATGPT_CAPTURE_PAYLOAD", + payload: fallbackPayload + }); + } + } + })(); +})(); diff --git a/applications/chatgpt-capture/src/connectors/chatgpt/chatgptConnector.js b/applications/chatgpt-capture/src/connectors/chatgpt/chatgptConnector.js new file mode 100644 index 0000000..e0c9536 --- /dev/null +++ b/applications/chatgpt-capture/src/connectors/chatgpt/chatgptConnector.js @@ -0,0 +1,51 @@ +import { prepareChatGPTForCapture } from "./chatgptPreprocessor.js"; +import { discoverChatGPTTurns } from "./chatgptDiscovery.js"; +import { renderChatGPTTurnToMarkdown } from "./chatgptTurnRenderer.js"; +import { createConversationModel } from "../../core/conversationModel.js"; +import { serializeConversationToMarkdown } from "../../core/markdownSerializer.js"; + +export async function captureConversation(options = {}) { + const root = options.root || document; + const turns = discoverChatGPTTurns(root); + const preprocessed = await prepareChatGPTForCapture({ root, turns }); + const renderedTurns = turns.map((turn) => { + const markdown = renderChatGPTTurnToMarkdown(turn); + const captureStatus = markdown.includes("[Turn not rendered") ? "not_rendered" : "captured"; + return { + ...turn, + captureStatus, + contentType: "text", + markdown, + debug: { + rendered: captureStatus === "captured" + } + }; + }); + + const conversation = createConversationModel({ + sourcePlatform: "ChatGPT", + title: options.title || document.title || "ChatGPT Conversation", + url: options.url || window.location.href, + capturedAt: options.capturedAt || new Date().toISOString(), + warning: renderedTurns.some((turn) => turn.captureStatus === "not_rendered") + ? "Some conversation turns were not rendered in the DOM. Scroll through the conversation and capture again." + : undefined, + turns: renderedTurns, + debug: { + turnSectionsFound: renderedTurns.length, + renderedTurnsCaptured: renderedTurns.filter((turn) => turn.captureStatus === "captured").length, + notRenderedTurns: renderedTurns.filter((turn) => turn.captureStatus === "not_rendered").length, + userTurns: renderedTurns.filter((turn) => turn.role === "user").length, + assistantTurns: renderedTurns.filter((turn) => turn.role === "assistant").length, + unknownTurns: renderedTurns.filter((turn) => turn.role === "unknown").length + } + }); + + const markdown = serializeConversationToMarkdown(conversation); + return { + conversation, + markdown, + debug: conversation.debug, + preprocessed + }; +} diff --git a/applications/chatgpt-capture/src/connectors/chatgpt/chatgptDiscovery.js b/applications/chatgpt-capture/src/connectors/chatgpt/chatgptDiscovery.js new file mode 100644 index 0000000..1a4d249 --- /dev/null +++ b/applications/chatgpt-capture/src/connectors/chatgpt/chatgptDiscovery.js @@ -0,0 +1,49 @@ +const TURN_SECTION_SELECTOR = 'section[data-testid^="conversation-turn-"][data-turn]'; + +function normalizeRole(role) { + const value = (role || "unknown").toLowerCase(); + if (value === "user" || value === "you") { + return "user"; + } + if (value === "assistant" || value === "chatgpt") { + return "assistant"; + } + if (value === "system") { + return "system"; + } + if (value === "tool") { + return "tool"; + } + return "unknown"; +} + +function parseTurnNumber(section) { + const match = (section?.getAttribute("data-testid") || "").match(/conversation-turn-(\d+)/i); + return match ? Number(match[1]) : NaN; +} + +export function discoverChatGPTTurns(root = document) { + const sections = Array.from(root.querySelectorAll(TURN_SECTION_SELECTOR)); + const orderedSections = sections + .map((section, index) => ({ + section, + sortIndex: Number.isNaN(parseTurnNumber(section)) ? index : parseTurnNumber(section) + })) + .sort((left, right) => left.sortIndex - right.sortIndex); + + return orderedSections.map((entry, index) => { + const section = entry.section; + const explicitRole = normalizeRole(section?.getAttribute("data-turn") || ""); + const authorRoleNode = section?.querySelector('[data-message-author-role]'); + const fallbackRole = authorRoleNode ? normalizeRole(authorRoleNode.getAttribute("data-message-author-role") || "") : "unknown"; + const role = explicitRole !== "unknown" ? explicitRole : fallbackRole; + + return { + turnIndex: index, + turnId: section?.getAttribute("data-testid") || `conversation-turn-${index + 1}`, + role, + sourceElement: section, + sourcePlatform: "ChatGPT" + }; + }); +} diff --git a/applications/chatgpt-capture/src/connectors/chatgpt/chatgptPreprocessor.js b/applications/chatgpt-capture/src/connectors/chatgpt/chatgptPreprocessor.js new file mode 100644 index 0000000..a1406ab --- /dev/null +++ b/applications/chatgpt-capture/src/connectors/chatgpt/chatgptPreprocessor.js @@ -0,0 +1,16 @@ +export async function prepareChatGPTForCapture(options = {}) { + const turns = options.turns || []; + if (turns.length > 0 && typeof options.root?.querySelectorAll === "function") { + turns.forEach((turn) => { + if (turn?.sourceElement?.scrollIntoView) { + turn.sourceElement.scrollIntoView({ block: "center", inline: "nearest" }); + } + }); + } + + return { + prepared: true, + message: "Preparing conversation for capture...", + scrolledTurns: turns.length + }; +} diff --git a/applications/chatgpt-capture/src/connectors/chatgpt/chatgptTurnRenderer.js b/applications/chatgpt-capture/src/connectors/chatgpt/chatgptTurnRenderer.js new file mode 100644 index 0000000..4983286 --- /dev/null +++ b/applications/chatgpt-capture/src/connectors/chatgpt/chatgptTurnRenderer.js @@ -0,0 +1,79 @@ +import { htmlToMarkdown } from "../../core/htmlToMarkdown.js"; + +function normalizeWhitespace(text) { + if (!text) { + return ""; + } + + return text + .replace(/\u00a0/g, " ") + .replace(/[ \t]+\n/g, "\n") + .replace(/\n{3,}/g, "\n\n") + .replace(/[ \t]{2,}/g, " ") + .trim(); +} + +function getTurnContentRoot(turn) { + const section = turn?.sourceElement; + if (!section) { + return null; + } + + if (turn.role === "assistant") { + const assistantMessage = section.querySelector('[data-message-author-role="assistant"]'); + return assistantMessage?.querySelector(".markdown") || assistantMessage || section.querySelector(".markdown") || section; + } + + if (turn.role === "user") { + const userMessage = section.querySelector('[data-message-author-role="user"]'); + return userMessage || section; + } + + return section; +} + +function isRenderCandidate(node) { + if (!node || node.nodeType !== Node.ELEMENT_NODE) { + return false; + } + + if (node.matches('button, [role="button"], [aria-hidden="true"], [hidden], .hidden, .sr-only, nav, aside, header, footer, form, input, textarea, select, svg, script, style')) { + return false; + } + + return true; +} + +export function renderChatGPTTurnToMarkdown(turn) { + const section = turn?.sourceElement; + const role = turn?.role || "unknown"; + const contentRoot = getTurnContentRoot(turn); + const placeholder = "[Turn not rendered in DOM. Scroll this turn into view and recapture.]"; + + if (!section || !contentRoot) { + return placeholder; + } + + const rect = typeof section.getBoundingClientRect === "function" ? section.getBoundingClientRect() : null; + const isVirtualized = section.getAttribute("data-is-intersecting") === "false"; + const hasRenderableContent = Boolean(contentRoot && Array.from(contentRoot.querySelectorAll("p, pre, ul, ol, li, blockquote, h1, h2, h3, h4, h5, h6, table, img, [data-message-author-role]")) + .filter(isRenderCandidate) + .some((node) => (node.textContent || "").trim().length > 0)); + + if (isVirtualized || (rect && rect.height <= 0 && rect.width <= 0) || (!hasRenderableContent && !(contentRoot.textContent || "").trim())) { + return placeholder; + } + + const clone = contentRoot.cloneNode(true); + clone.querySelectorAll('button, [role="button"], [aria-hidden="true"], [hidden], .hidden, .sr-only, nav, aside, header, footer, form, input, textarea, select, svg, script, style').forEach((child) => child.remove()); + + if (role === "user") { + const userMessage = clone.querySelector('[data-message-author-role="user"]') || clone; + const markdown = htmlToMarkdown(userMessage, { role }); + return normalizeWhitespace(markdown || placeholder); + } + + const assistantMessage = clone.querySelector('[data-message-author-role="assistant"]') || clone.querySelector(".markdown") || clone; + const markdown = htmlToMarkdown(assistantMessage, { role }); + return normalizeWhitespace(markdown || placeholder); +} diff --git a/applications/chatgpt-capture/src/core/conversationModel.js b/applications/chatgpt-capture/src/core/conversationModel.js new file mode 100644 index 0000000..f735524 --- /dev/null +++ b/applications/chatgpt-capture/src/core/conversationModel.js @@ -0,0 +1,32 @@ +export function createConversationTurn(turn) { + return { + turnIndex: turn?.turnIndex ?? 0, + turnId: turn?.turnId || "", + role: turn?.role || "unknown", + captureStatus: turn?.captureStatus || "captured", + sourcePlatform: turn?.sourcePlatform || "ChatGPT", + contentType: turn?.contentType || "text", + sourceElement: turn?.sourceElement || null, + markdown: turn?.markdown || "", + debug: turn?.debug || {} + }; +} + +export function createConversationModel(payload = {}) { + return { + sourcePlatform: payload?.sourcePlatform || "ChatGPT", + title: payload?.title || "ChatGPT Conversation", + url: payload?.url || "", + capturedAt: payload?.capturedAt || new Date().toISOString(), + warning: payload?.warning, + turns: (payload?.turns || []).map(createConversationTurn), + debug: { + turnSectionsFound: payload?.debug?.turnSectionsFound || 0, + renderedTurnsCaptured: payload?.debug?.renderedTurnsCaptured || 0, + notRenderedTurns: payload?.debug?.notRenderedTurns || 0, + userTurns: payload?.debug?.userTurns || 0, + assistantTurns: payload?.debug?.assistantTurns || 0, + unknownTurns: payload?.debug?.unknownTurns || 0 + } + }; +} diff --git a/applications/chatgpt-capture/src/core/htmlToMarkdown.js b/applications/chatgpt-capture/src/core/htmlToMarkdown.js new file mode 100644 index 0000000..eaf24c8 --- /dev/null +++ b/applications/chatgpt-capture/src/core/htmlToMarkdown.js @@ -0,0 +1,232 @@ +function normalizeWhitespace(text) { + if (!text) { + return ""; + } + + return text + .replace(/\u00a0/g, " ") + .replace(/[ \t]+\n/g, "\n") + .replace(/\n{3,}/g, "\n\n") + .replace(/[ \t]{2,}/g, " ") + .trim(); +} + +function normalizeInlineText(text) { + return (text || "") + .replace(/\u00a0/g, " ") + .replace(/\s+/g, " ") + .replace(/\s+([,.;:!?])/g, "$1") + .trim(); +} + +function escapeMarkdownLinkText(text) { + return String(text || "") + .replace(/\[/g, "\\[") + .replace(/\]/g, "\\]") + .trim(); +} + +function escapeMarkdownLinkUrl(url) { + return String(url || "") + .replace(/\)/g, "%29") + .trim(); +} + +function isSafeHref(href) { + const value = (href || "").trim(); + if (!value || /^javascript:/i.test(value) || /^data:/i.test(value) || /^blob:/i.test(value)) { + return false; + } + return /^https?:\/\//i.test(value) || /^mailto:/i.test(value); +} + +function isInlineElement(node) { + return Boolean( + node && + node.nodeType === Node.ELEMENT_NODE && + node.matches("strong, b, em, i, code, span, a, small, sub, sup, u, s") + ); +} + +function getCodeLanguage(node) { + const className = (node?.className || "").toString(); + const languageMatch = className.match(/language-([a-z0-9_+-]+)/i) || className.match(/lang-([a-z0-9_+-]+)/i); + const dataLanguage = node?.getAttribute?.("data-language") || node?.getAttribute?.("data-lang"); + return languageMatch ? languageMatch[1] : dataLanguage || ""; +} + +function anchorToMarkdown(anchorNode, role) { + const href = anchorNode.getAttribute("href") || ""; + const label = normalizeInlineText(childrenToMarkdown(anchorNode, role)) || normalizeInlineText(anchorNode.textContent || "") || href; + if (!isSafeHref(href)) { + return label; + } + return `[${escapeMarkdownLinkText(label)}](${escapeMarkdownLinkUrl(href)})`; +} + +function imageToMarkdown(img, role) { + const alt = normalizeInlineText(img.getAttribute("alt") || ""); + const src = img.getAttribute("src") || ""; + const label = alt || src || "image"; + if (role === "user") { + return `[Attached image: ${label}](${src || ""})`; + } + if (src) { + return `![Generated image: ${label}](${src})`; + } + return `![Generated image: ${label}]`; +} + +function listToMarkdown(listNode, role, depth = 0) { + const listType = listNode.tagName.toLowerCase(); + const items = Array.from(listNode.children || []).filter((child) => child.matches("li")); + const prefix = depth > 0 ? " ".repeat(depth) : ""; + + return items + .map((item, index) => { + const marker = listType === "ol" ? `${index + 1}. ` : "- "; + const inlineContent = childrenToMarkdown(item, role).trim(); + const nestedItems = Array.from(item.children || []).filter((child) => child.matches("ul, ol")); + const nestedMarkdown = nestedItems.map((nestedList) => listToMarkdown(nestedList, role, depth + 1)).filter(Boolean).join("\n"); + const content = inlineContent || normalizeInlineText(item.textContent || ""); + return [ + `${prefix}${marker}${content}`, + nestedMarkdown ? `${prefix} ${nestedMarkdown.replace(/\n/g, `\n${prefix} `)}` : "" + ].filter(Boolean).join("\n"); + }) + .filter(Boolean) + .join("\n"); +} + +function tableToMarkdown(tableNode, role) { + const rows = Array.from(tableNode.querySelectorAll("tr")); + if (rows.length === 0) { + return ""; + } + const headerRow = rows.find((row) => row.querySelector("th")) || rows[0]; + const bodyRows = rows.filter((row) => row !== headerRow); + const headerCells = Array.from(headerRow.querySelectorAll("th, td")) + .map((cell) => normalizeInlineText(childrenToMarkdown(cell, role))) + .filter(Boolean); + if (headerCells.length === 0) { + return ""; + } + const body = bodyRows.map((row) => { + const cells = Array.from(row.querySelectorAll("td")) + .map((cell) => normalizeInlineText(childrenToMarkdown(cell, role))) + .filter(Boolean); + return cells.length > 0 ? `| ${cells.join(" | ")} |` : ""; + }).filter(Boolean); + const headerLine = `| ${headerCells.join(" | ")} |`; + const separatorLine = `| ${headerCells.map(() => "---").join(" | ")} |`; + return [headerLine, separatorLine, ...body].join("\n"); +} + +function codeBlockToMarkdown(preNode) { + const codeNode = preNode.querySelector("code") || preNode; + const language = getCodeLanguage(codeNode); + const content = (codeNode.textContent || "") + .replace(/\r\n/g, "\n") + .replace(/\n{3,}/g, "\n\n") + .trim(); + return [language ? `\`\`\`${language}` : "```", content, "```"].filter(Boolean).join("\n"); +} + +function blockquoteToMarkdown(blockquoteNode, role) { + return Array.from(blockquoteNode.childNodes) + .map((child) => { + if (child.nodeType === Node.TEXT_NODE) { + return `> ${normalizeInlineText(child.textContent || "")}`; + } + if (child.nodeType !== Node.ELEMENT_NODE) { + return ""; + } + const markdown = nodeToMarkdown(child, role); + return markdown.split(/\n/).map((line) => (line ? `> ${line}` : ">")).join("\n"); + }) + .filter(Boolean) + .join("\n"); +} + +function childrenToMarkdown(node, role) { + const parts = []; + Array.from(node.childNodes).forEach((child) => { + if (child.nodeType === Node.TEXT_NODE) { + const text = normalizeInlineText(child.textContent || ""); + if (text) { + parts.push(text); + } + return; + } + if (child.nodeType !== Node.ELEMENT_NODE) { + return; + } + if (child.matches("strong, b")) { + parts.push(`**${childrenToMarkdown(child, role)}**`); + } else if (child.matches("em, i")) { + parts.push(`*${childrenToMarkdown(child, role)}*`); + } else if (child.matches("code")) { + parts.push(`\`${normalizeInlineText(child.textContent || "")}\``); + } else if (child.matches("a")) { + parts.push(anchorToMarkdown(child, role)); + } else if (child.matches("img")) { + parts.push(imageToMarkdown(child, role)); + } else { + parts.push(nodeToMarkdown(child, role)); + } + }); + return parts.join(" ").trim(); +} + +function nodeToMarkdown(node, role) { + if (!node) { + return ""; + } + if (node.nodeType === Node.TEXT_NODE) { + return normalizeInlineText(node.textContent || ""); + } + if (node.nodeType !== Node.ELEMENT_NODE) { + return ""; + } + if (node.matches("pre")) { + return codeBlockToMarkdown(node); + } + if (node.matches("code")) { + return `\`${normalizeInlineText(node.textContent || "")}\``; + } + if (node.matches("br")) { + return "\n"; + } + if (node.matches("blockquote")) { + return blockquoteToMarkdown(node, role); + } + if (node.matches("table")) { + return tableToMarkdown(node, role); + } + if (node.matches("ul, ol")) { + return listToMarkdown(node, role); + } + if (node.matches("h1, h2, h3, h4, h5, h6")) { + const level = parseInt(node.tagName.charAt(1), 10) || 1; + return `${"#".repeat(level)} ${childrenToMarkdown(node, role)}`; + } + if (node.matches("p")) { + return childrenToMarkdown(node, role); + } + if (isInlineElement(node)) { + return childrenToMarkdown(node, role); + } + return childrenToMarkdown(node, role); +} + +function cleanupWhitespace(markdown) { + return normalizeWhitespace(markdown || "").replace(/\n{3,}/g, "\n\n"); +} + +export function htmlToMarkdown(root, options = {}) { + const role = options.role || "unknown"; + if (!root) { + return ""; + } + return cleanupWhitespace(nodeToMarkdown(root, role)); +} diff --git a/applications/chatgpt-capture/src/core/markdownSerializer.js b/applications/chatgpt-capture/src/core/markdownSerializer.js new file mode 100644 index 0000000..63d40c8 --- /dev/null +++ b/applications/chatgpt-capture/src/core/markdownSerializer.js @@ -0,0 +1,52 @@ +function normalizeRoleHeading(role) { + switch ((role || "unknown").toLowerCase()) { + case "user": + return "## User"; + case "assistant": + return "## Assistant"; + case "system": + return "## System"; + case "tool": + return "## Tool"; + default: + return "## Unknown"; + } +} + +function normalizeMessageContent(content) { + return (content || "").trim(); +} + +export function serializeConversationToMarkdown(conversation) { + const lines = []; + lines.push(`# ${conversation?.title || "ChatGPT Conversation"}`); + lines.push(""); + if (conversation?.capturedAt) { + lines.push(`Captured: ${conversation.capturedAt}`); + } + lines.push(`Source Platform: ${conversation?.sourcePlatform || "ChatGPT"}`); + if (conversation?.url) { + lines.push(`Source URL: ${conversation.url}`); + } + if (conversation?.warning) { + lines.push(""); + lines.push(`Warning: ${conversation.warning}`); + } + lines.push(""); + lines.push("---"); + lines.push(""); + + const turns = Array.isArray(conversation?.turns) ? conversation.turns : []; + turns.forEach((turn, index) => { + lines.push(normalizeRoleHeading(turn?.role)); + lines.push(""); + lines.push(normalizeMessageContent(turn?.markdown)); + if (index < turns.length - 1) { + lines.push(""); + lines.push("---"); + lines.push(""); + } + }); + + return `${lines.join("\n").trim()}\n`; +} diff --git a/applications/chatgpt-capture/src/filename.js b/applications/chatgpt-capture/src/filename.js new file mode 100644 index 0000000..5262104 --- /dev/null +++ b/applications/chatgpt-capture/src/filename.js @@ -0,0 +1,44 @@ +function formatDateForFilename(dateValue) { + const date = dateValue ? new Date(dateValue) : new Date(); + + if (Number.isNaN(date.getTime())) { + return new Date().toISOString().slice(0, 10); + } + + return date.toISOString().slice(0, 10); +} + +function sanitizeFilenamePart(value) { + return String(value || "") + .normalize("NFKD") + .replace(/[\u0000-\u001F\u007F]/g, "") + .replace(/[<>:"/\\|?*]/g, " ") + .replace(/\s+/g, " ") + .replace(/^\.+|\.+$/g, "") + .replace(/^\s+|\s+$/g, "") + .trim(); +} + +function truncateFilename(filename, maxLength) { + if (!filename) { + return filename; + } + + if (filename.length <= maxLength) { + return filename; + } + + const extension = filename.slice(filename.lastIndexOf(".")); + const base = filename.slice(0, filename.lastIndexOf(".")); + const safeExtension = extension || ".md"; + const safeBase = base.slice(0, Math.max(1, maxLength - safeExtension.length - 1)).trimEnd(); + return `${safeBase}${safeExtension}`; +} + +export function createConversationFilename(capturePayload) { + const date = formatDateForFilename(capturePayload?.capturedAt); + const platform = sanitizeFilenamePart(capturePayload?.sourcePlatform || "ChatGPT") || "ChatGPT"; + const title = sanitizeFilenamePart(capturePayload?.title || "Conversation") || "Conversation"; + const baseFilename = `${date} - ${platform} - ${title}.md`; + return truncateFilename(baseFilename, 180); +} diff --git a/applications/chatgpt-capture/src/markdownNormalizer.js b/applications/chatgpt-capture/src/markdownNormalizer.js new file mode 100644 index 0000000..c8cd8e3 --- /dev/null +++ b/applications/chatgpt-capture/src/markdownNormalizer.js @@ -0,0 +1,120 @@ +function normalizeTitle(title) { + const normalized = (title || "ChatGPT Conversation") + .replace(/\s+/g, " ") + .trim(); + + return normalized || "ChatGPT Conversation"; +} + +function normalizeRoleHeading(role) { + switch ((role || "unknown").toLowerCase()) { + case "user": + return "## User"; + case "assistant": + return "## Assistant"; + case "system": + return "## System"; + default: + return "## Unknown"; + } +} + +function collapseExcessBlankLinesOutsideCodeBlocks(text) { + const lines = (text || "").split(/\r?\n/); + const result = []; + let inCodeBlock = false; + + for (const line of lines) { + const trimmed = line.trim(); + const fence = trimmed.startsWith("```"); + + if (fence) { + if (!inCodeBlock) { + inCodeBlock = true; + } else { + inCodeBlock = false; + } + result.push(line); + continue; + } + + if (inCodeBlock) { + result.push(line); + continue; + } + + if (trimmed === "") { + if (result.length === 0 || result[result.length - 1].trim() !== "") { + result.push(""); + } + continue; + } + + result.push(line); + } + + return result.join("\n").trim(); +} + +function normalizeMessageContent(content) { + return collapseExcessBlankLinesOutsideCodeBlocks((content || "").trim()); +} + +export function normalizeConversationToMarkdown(capturePayload) { + const payload = capturePayload || {}; + const title = normalizeTitle(payload.title); + const capturedAt = payload.capturedAt || ""; + const sourcePlatform = payload.sourcePlatform || "Unknown"; + const sourceUrl = payload.url || ""; + const messages = Array.isArray(payload.messages) ? payload.messages : []; + + const lines = []; + lines.push(`# ${title}`); + lines.push(""); + + if (capturedAt) { + lines.push(`Captured: ${capturedAt}`); + } + + lines.push(`Source Platform: ${sourcePlatform}`); + if (sourceUrl) { + lines.push(`Source URL: ${sourceUrl}`); + } + + lines.push(""); + lines.push("---"); + lines.push(""); + + if (payload.error) { + lines.push("## Capture Error"); + lines.push(""); + lines.push(payload.error); + lines.push(""); + lines.push("---"); + lines.push(""); + } + + if (messages.length === 0) { + lines.push("## Capture Notice"); + lines.push(""); + lines.push("No conversation messages were detected."); + lines.push(""); + lines.push("---"); + lines.push(""); + } else { + messages.forEach((message, index) => { + lines.push(normalizeRoleHeading(message?.role)); + lines.push(""); + lines.push(normalizeMessageContent(message?.content)); + if (index < messages.length - 1) { + lines.push(""); + lines.push("---"); + lines.push(""); + } + }); + } + + return `${lines.join("\n").trim()}\n`; +} + +globalThis.normalizeConversationToMarkdown = normalizeConversationToMarkdown; diff --git a/codex/chatgpt-capture-mvp/README.md b/codex/chatgpt-capture-mvp/README.md new file mode 100644 index 0000000..e69de29 diff --git a/codex/chatgpt-capture-mvp/STATUS.md b/codex/chatgpt-capture-mvp/STATUS.md new file mode 100644 index 0000000..2b75db3 --- /dev/null +++ b/codex/chatgpt-capture-mvp/STATUS.md @@ -0,0 +1,25 @@ +# ChatGPT Capture MVP + +Current Phase + +Implementation + +Current Task + +0003c-render-all-turns + +Completed + +✓ 0001 +✓ 0002 +✓ 0003 +✓ 0003a +✓ 0003b + +Blocked + +None + +Next + +0004-markdown-normalizer \ No newline at end of file diff --git a/codex/chatgpt-capture-mvp/debugging/0001-formatting-defects.md b/codex/chatgpt-capture-mvp/debugging/0001-formatting-defects.md new file mode 100644 index 0000000..25b944b --- /dev/null +++ b/codex/chatgpt-capture-mvp/debugging/0001-formatting-defects.md @@ -0,0 +1,86 @@ +Fix ChatGPT DOM extractor formatting defects discovered during manual testing. + +Observed defects: +1. Bold text is being extracted with comma-like artifacts instead of Markdown emphasis. + +Example bad output: +what we know +, +what is plausible +, and +what is not supported by the evidence +. + +Expected output: +**what we know**, **what is plausible**, and **what is not supported by the evidence**. + +2. Quoted text is being surrounded by em-dash artifacts. + +Example bad output: +So the author's central educational point— +"Jet streams matter." +—is absolutely correct. + +Expected output: +So the author's central educational point — "Jet streams matter." — is absolutely correct. + +Likely cause: +The extractor is walking DOM nodes and treating presentational or punctuation-only nodes incorrectly, especially around strong/bold tags and quotation/punctuation boundaries. + +Requirements: +1. Update src/chatgptExtractor.js only unless integration requires a small change elsewhere. + +2. Improve inline formatting extraction: + - Convert and to Markdown bold: **text** + - Convert and to Markdown italics: *text* + - Convert inline to `text` + - Preserve surrounding punctuation in natural sentence order. + - Do not emit isolated punctuation artifacts on separate lines. + +3. Fix punctuation spacing: + - Do not place commas, periods, semicolons, colons, exclamation marks, or question marks on their own line. + - Attach closing punctuation to the preceding token. + - Preserve spaces between words. + - Avoid inserting line breaks between inline elements unless the source node is block-level. + +4. Fix quote/em-dash handling: + - Do not treat em dashes as block separators. + - Keep em dashes inline with surrounding text. + - Normalize this pattern: + text— + "quote" + —text + into readable inline prose where possible. + +5. Preserve block formatting: + - Paragraphs should remain paragraphs. + - Lists should remain readable. + - Code blocks should still use fenced Markdown. + - Headings should remain Markdown headings where already implemented. + +6. Add helper functions if useful: + - isInlineElement(node) + - isBlockElement(node) + - appendInlineText(parts, text) + - normalizeInlineMarkdown(text) + - cleanupPunctuationSpacing(text) + +7. Add lightweight manual test fixtures or comments with examples: + - bold phrase followed by comma + - bold phrase followed by comma and another bold phrase + - quoted phrase surrounded by em dashes + - assistant message with paragraph + list + code block + +8. Do not add AI processing. +9. Do not change Markdown normalizer behavior unless absolutely necessary. +10. Do not change download writer behavior. + +Acceptance criteria: +- Bold text exports as **bold text**. +- Italics export as *italic text*. +- Inline code exports as `inline code`. +- Commas remain attached to the correct phrase. +- Em dashes remain inline and readable. +- No punctuation-only lines are introduced by inline formatting. +- Code blocks remain intact. +- Message order and role detection still work. \ No newline at end of file diff --git a/codex/chatgpt-capture-mvp/debugging/0002-block-defects.md b/codex/chatgpt-capture-mvp/debugging/0002-block-defects.md new file mode 100644 index 0000000..252aa33 --- /dev/null +++ b/codex/chatgpt-capture-mvp/debugging/0002-block-defects.md @@ -0,0 +1,100 @@ +Fix ChatGPT DOM extractor block-structure defects discovered during manual testing. + +Observed defect: +Large assistant responses are being flattened into plain lines. Paragraph breaks, list structure, indentation, and inline emphasis boundaries are being lost. + +Example bad output: +There's a real pattern here, but it's useful to separate +public narratives +from +the underlying issues +. +Over the past several decades... + +Expected behavior: +There's a real pattern here, but it's useful to separate **public narratives** from **the underlying issues**. + +Over the past several decades, there have indeed been many highly publicized predictions of catastrophic outcomes: + +- Nuclear war during the Cold War. +- Concerns about global cooling in the 1970s... +- Y2K causing widespread infrastructure failures. + +Root cause: +The extractor is likely using innerText/textContent or recursively joining nodes without distinguishing inline elements from block elements. + +Requirements: +1. Update src/chatgptExtractor.js. + +2. Implement DOM-to-Markdown extraction that treats block and inline elements differently: + - Inline elements must remain inline. + - Block elements must create paragraph/list/code boundaries. + - Do not insert line breaks around inline tags like strong, em, span, a, code. + +3. Preserve paragraphs: + - Each

becomes one paragraph. + - Paragraphs separated by one blank line. + - Inline formatting inside paragraphs must remain inline. + +4. Preserve unordered lists: + -

  • Item
becomes: + - Item + - Preserve nested list indentation where practical. + +5. Preserve ordered lists: + -
  1. Item
becomes: + 1. Item + 2. Item + +6. Preserve blockquotes where detectable: + - Prefix quoted block lines with >. + +7. Preserve headings: + - h1-h6 become Markdown headings. + +8. Preserve code blocks: + -
...
becomes fenced code. + - Do not wrap or reformat code content. + +9. Preserve inline formatting: + - / → **text** + - / → *text* + - inline → `text` + - may preserve readable text only for MVP. + +10. Add a cleaner that operates after Markdown generation: + - Remove punctuation-only lines caused by extraction. + - Collapse 3+ blank lines to 2. + - Remove spaces before punctuation. + - Preserve content inside fenced code blocks unchanged. + +11. Do not use innerText for whole message containers as the primary extraction method. + - It may be used only as a fallback when DOM parsing fails. + +12. Add small local test fixtures or comments for: + - paragraph with bold inline phrases + - paragraph followed by unordered list + - paragraph followed by numbered list + - nested list if easy + - code block + - quote/em dash sentence + +Suggested helper functions: +- nodeToMarkdown(node, context) +- childrenToInlineMarkdown(node, context) +- blockChildrenToMarkdown(node, context) +- listToMarkdown(listNode, context) +- listItemToMarkdown(liNode, context) +- codeBlockToMarkdown(preNode) +- cleanupMarkdownOutsideCodeFences(markdown) + +Acceptance criteria: +- The sample assistant response preserves paragraphs. +- Catastrophe examples become a Markdown bullet list. +- The “important distinctions” examples become separate paragraphs or list items if ChatGPT rendered them as list items. +- Bold phrases remain inline, not on separate lines. +- Periods and commas remain attached to the correct sentence. +- Code blocks remain intact. +- Message order and role detection are unchanged. +- No LLM calls. +- No Project Thoth application dependency. \ No newline at end of file diff --git a/codex/chatgpt-capture-mvp/debugging/0003-table-defects.md b/codex/chatgpt-capture-mvp/debugging/0003-table-defects.md new file mode 100644 index 0000000..71757c3 --- /dev/null +++ b/codex/chatgpt-capture-mvp/debugging/0003-table-defects.md @@ -0,0 +1,76 @@ +Fix ChatGPT DOM extractor table handling. + +Observed defect: +Tables in ChatGPT responses are being flattened into line-by-line text instead of Markdown tables. + +Example bad output: +Goal + +Good Plant Choices + +Attract deer + +Hostas, daylilies, many ornamentals + +Feed pollinators + +Native flowering perennials, shrubs, trees + +Expected output: +| Goal | Good Plant Choices | +|---|---| +| Attract deer | Hostas, daylilies, many ornamentals | +| Feed pollinators | Native flowering perennials, shrubs, trees | +| Prevent deer damage | Deer-resistant species | +| Support overall biodiversity | A mix of native plants with different bloom times | + +Requirements: +1. Update src/chatgptExtractor.js. + +2. Add table-to-Markdown support: + - Detect elements. + - Extract rows from . + - Extract header cells from , use that as the header. + - If no exists but the first row uses
. + - Extract data cells from . + - Preserve cell order. + - Trim cell text. + - Convert inline formatting inside cells using existing inline extractor. + +3. Markdown table rules: + - First row should become the header row. + - If the table has
, use that row as the header. + - If no header exists, use the first row as the header for MVP. + - Add separator row using `---`. + - Escape pipe characters inside cells as `\|`. + - Replace internal newlines in cells with `
` or a single space. + +4. Example output format: + | Goal | Good Plant Choices | + |---|---| + | Attract deer | Hostas, daylilies, many ornamentals | + +5. Ensure table extraction is treated as a block element: + - Tables should be separated from surrounding paragraphs by blank lines. + - Do not also extract duplicate flattened table text from child nodes. + +6. Preserve other existing behavior: + - Paragraphs remain paragraphs. + - Lists remain Markdown lists. + - Code blocks remain fenced code. + - Bold/italic/inline code remain inline. + - Message order and role detection unchanged. + +Suggested helper: +function tableToMarkdown(tableNode, context) { + // return markdown table string +} + +Acceptance criteria: +- ChatGPT tables export as valid Markdown tables. +- Headers and rows are preserved. +- Pipe characters inside cells are escaped. +- Tables are not duplicated as flattened text. +- Surrounding prose remains intact. +- No LLM calls. +- No Project Thoth application dependency. \ No newline at end of file diff --git a/codex/chatgpt-capture-mvp/debugging/0004-external-link-defect.md b/codex/chatgpt-capture-mvp/debugging/0004-external-link-defect.md new file mode 100644 index 0000000..4338e0d --- /dev/null +++ b/codex/chatgpt-capture-mvp/debugging/0004-external-link-defect.md @@ -0,0 +1,68 @@ +Fix ChatGPT DOM extractor external link handling. + +Observed defect: +External links in ChatGPT responses are being extracted as plain text only. + +Example bad output: +Services such as: + +Amazon Prime Video +Apple TV +Fandango at Home (Vudu) + +Expected behavior: +Preserve links as Markdown links when href is available: + +- [Amazon Prime Video](https://...) +- [Apple TV](https://...) +- [Fandango at Home (Vudu)](https://...) + +Requirements: +1. Update src/chatgptExtractor.js. + +2. Add Markdown link support for elements: + - Extract visible link text using the existing inline extractor. + - Extract href. + - If href is valid, output `[text](href)`. + - If href is missing or unsafe, fall back to plain visible text. + +3. Preserve surrounding inline formatting: + - Links inside paragraphs remain inline. + - Links inside lists remain list items. + - Links inside table cells remain Markdown links. + - Do not force every link onto its own line unless the original block structure does. + +4. URL handling: + - Preserve http:// and https:// links. + - Preserve mailto: links if encountered. + - Ignore or flatten javascript:, data:, blob:, and empty href values. + - Decode or normalize only as needed for readable Markdown. + - Escape closing parentheses in URLs if necessary. + +5. Link text handling: + - Trim excessive whitespace. + - If link text is empty, use the href as the visible text. + - Escape square brackets in link text. + - Do not duplicate the URL if ChatGPT already displays the raw URL as the link text. + +6. Preserve existing behavior: + - Paragraphs remain paragraphs. + - Lists remain Markdown lists. + - Tables remain Markdown tables. + - Code blocks remain fenced code. + - Bold, italics, and inline code remain inline. + - Message order and role detection unchanged. + +Suggested helper functions: +- linkToMarkdown(anchorNode, context) +- isSafeHref(href) +- escapeMarkdownLinkText(text) +- escapeMarkdownLinkUrl(url) + +Acceptance criteria: +- External links export as valid Markdown links. +- Plain text without links remains unchanged. +- Unsafe href values are not preserved as Markdown links. +- Links in lists, paragraphs, and tables remain readable. +- No LLM calls. +- No Project Thoth application dependency. \ No newline at end of file diff --git a/codex/chatgpt-capture-mvp/debugging/0005-implementation3a-over-capture-defect.md b/codex/chatgpt-capture-mvp/debugging/0005-implementation3a-over-capture-defect.md new file mode 100644 index 0000000..a8c177a --- /dev/null +++ b/codex/chatgpt-capture-mvp/debugging/0005-implementation3a-over-capture-defect.md @@ -0,0 +1,47 @@ +Fix Task 3A conversation discovery over-capture and duplication. + +Problem: +The extractor is now capturing the ChatGPT application shell, sidebar, navigation, composer UI, and duplicate parent/child containers. It also captures large flattened blocks containing both "You said" and "ChatGPT said" instead of discrete messages. + +Goal: +Discover only the conversation transcript stream and return one message object per actual conversation turn. + +Requirements: +1. Exclude ChatGPT app shell: + - Skip sidebar/nav/history/project/account areas. + - Skip composer/input area. + - Skip buttons, menus, share controls, footer text, scripts, and hidden UI. + +2. Scope discovery to the main conversation region: + - Prefer main[role="main"] or the primary conversation scroll container. + - Do not search the entire document body unless used only as a fallback. + +3. Deduplicate containers: + - Do not capture both a parent and its child if both contain the same message text. + - Prefer the smallest stable container that represents one message. + - Track DOM containment and normalized text fingerprints to remove duplicates. + +4. Split transcript turns: + - Do not allow one captured block to contain multiple "You said:" / "ChatGPT said:" turns. + - Each actual user or assistant turn should become a separate message. + +5. Keep fallback behavior, but only after filtering: + - Unknown containers are allowed, but they must come from the conversation region, not the app shell. + - Unknown should mean unknown conversation content, not random ChatGPT UI. + +6. Add debug counts: + - candidateContainersFound + - afterShellFilter + - afterDeduplication + - messagesReturned + - skippedShellContainers + - skippedDuplicateContainers + +Acceptance criteria: +- No "Skip to content" block. +- No sidebar chat history. +- No repeated full transcript blocks. +- No window.__oai_logHTML or script text. +- Actual user/assistant turns are separate messages. +- Image generation/edit placeholders may remain as unknown or unsupported, but only once. +- Completeness remains favored, but not by capturing the whole app shell. \ No newline at end of file diff --git a/codex/chatgpt-capture-mvp/debugging/0006-message-granularity-defect.md b/codex/chatgpt-capture-mvp/debugging/0006-message-granularity-defect.md new file mode 100644 index 0000000..c143b25 --- /dev/null +++ b/codex/chatgpt-capture-mvp/debugging/0006-message-granularity-defect.md @@ -0,0 +1,47 @@ +Fix ChatGPT extractor message granularity and role inheritance. + +Problem: +The extractor is now discovering conversation content, but it is treating paragraphs, list items, links, inline fragments, and code blocks as separate messages. This creates many Unknown blocks and misclassified User/Assistant fragments. + +Goal: +Return one message object per actual ChatGPT turn, not one message per child DOM block. + +Requirements: +1. Identify the outer message turn container first. + - A single user prompt should produce one User message. + - A single assistant response should produce one Assistant message. + - Child paragraphs, lists, tables, links, and code blocks must be content within that message, not separate messages. + +2. Role must be assigned from the outer message container. + - Child nodes must inherit the parent message role. + - Do not classify inline fragments independently. + - Do not classify text fragments like "re", "can", "own", "have" as User messages. + +3. Deduplicate repeated turns. + - The same user prompt should not appear twice. + - Use normalized text fingerprinting after role assignment. + +4. Fix link serialization. + - A Markdown link must be emitted as one inline unit. + - Do not split `[MakeMKV](https://www.makemkv.com)` into separate Unknown/Assistant blocks. + +5. Keep unsupported placeholders inside the parent turn. + - `[Unsupported ChatGPT content: ...]` should not be split across Unknown and Assistant blocks. + +6. Debug output should include: + - turnContainersFound + - messagesReturned + - duplicateTurnsSkipped + - orphanChildBlocksSkipped + - unknownTurnsReturned + +Acceptance criteria: +- The Jellyfin conversation starts with exactly one User message for the opening question. +- The first assistant answer is one Assistant message containing all options, lists, links, and code blocks. +- The Tiny Media Manager question is one User message. +- The SmartScreen answer is one Assistant message. +- The .nfo question is one User message. +- The .nfo answer is one Assistant message. +- No `.com)` standalone Assistant message. +- No `re`, `can`, `own`, `have`, or similar fragments as User messages. +- Unknown messages are allowed only for actual unsupported ChatGPT turn-level content, not paragraph/list/link fragments. \ No newline at end of file diff --git a/codex/chatgpt-capture-mvp/debugging/0007-resolve-connector-runtime.md b/codex/chatgpt-capture-mvp/debugging/0007-resolve-connector-runtime.md new file mode 100644 index 0000000..eb413ad --- /dev/null +++ b/codex/chatgpt-capture-mvp/debugging/0007-resolve-connector-runtime.md @@ -0,0 +1,142 @@ +# Work Order: Resolve connectorRuntime Module-Loading Regression + +## Current defect + +The browser extension no longer creates a Markdown file in Downloads. + +The extension console reports: + +```text +chrome-extension://.../connectorRuntime.js:1 +Uncaught SyntaxError: Cannot use import statement outside a module +``` + +A separate page-level error also appears: + +```text +QuotaExceededError: Failed to execute 'setItem' on 'Storage': +Setting the value of +'cache/user-JY5yPfTiyFSprhKcrYXMA2gG/9e03f04b-55ee-4773-bb7b-47df25276387/system-connectors' +exceeded the quota. +``` + +Treat the `QuotaExceededError` as originating from the ChatGPT application unless repository evidence shows that the extension interacts with that storage key. + +## Objective + +Restore the extension to the point where invoking capture reliably creates a Markdown file in the browser's Downloads location. + +## Required investigation + +1. Identify where `connectorRuntime.js` is loaded. +2. Identify the `import` statement or statements causing the parse failure. +3. Determine whether `connectorRuntime.js` is intended to be: + + * an ES module, + * a bundled traditional script, or + * imported by another module. +4. Compare the current loading configuration with the last working implementation or repository history, where available. +5. Determine whether the regression was introduced by: + + * adding an `import` statement to a non-module content script, + * changing the manifest entry, + * changing script injection logic, + * moving code without updating module boundaries, or + * failing to include the file in the build/bundle process. + +## Constraints + +* Make the smallest change necessary to restore the intended module-loading behavior. +* Do not redesign the DOM extractor. +* Do not change conversation-selection logic. +* Do not modify Markdown normalization except where directly required by the module-loading repair. +* Do not suppress the syntax error without addressing its cause. +* Do not treat the ChatGPT `QuotaExceededError` as the extension root cause without evidence. +* Preserve Manifest V3 compatibility. + +## Possible correction paths + +Evaluate the repository and choose the correction consistent with the existing architecture. Possible approaches include: + +### Module-based loading + +If `connectorRuntime.js` is intended to be an ES module, ensure it is loaded through a supported module mechanism. + +For an extension service worker, this may require: + +```json +{ + "background": { + "service_worker": "background.js", + "type": "module" + } +} +``` + +For a page or extension HTML document, this may require: + +```html + +``` + +Do not assume that declaring a content script as a module is supported in the same way. Verify how the file is being executed. + +### Bundled or non-module loading + +If `connectorRuntime.js` is a content script or injected traditional script, remove direct runtime `import` syntax by using the repository's intended bundling process or by restoring the prior dependency-loading pattern. + +Do not simply concatenate files unless that is already the project’s build design. + +### Dynamic import + +Use `import()` only if it is valid in the execution context and consistent with the extension architecture. Do not use it merely to hide a structural module problem. + +## Diagnostics + +Add or retain narrowly scoped diagnostic logging around: + +```text +[Thoth] connector runtime loaded +[Thoth] capture requested +[Thoth] extraction completed +[Thoth] markdown generated +[Thoth] download requested +[Thoth] download completed: +``` + +Errors should be logged with the failed stage and exception. + +## Acceptance criteria + +1. The extension loads without: + +```text +Cannot use import statement outside a module +``` + +2. Clicking the capture action reaches `connectorRuntime.js`. + +3. The extension creates a `.md` file in Downloads. + +4. The resulting file is non-empty. + +5. The extension console contains no uncaught extension-originated exception during capture. + +6. The ChatGPT page-level `QuotaExceededError`, if still present, does not prevent extension capture. + +7. Document: + + * the root cause, + * the file or manifest entry responsible, + * the corrective change, + * why the selected module-loading approach is correct, + * and the verification performed. + +## Out of scope + +* Improving DOM selectors +* Solving missing historical turns +* Improving table conversion +* Improving citation conversion +* Changing the conversation intermediate model +* General refactoring diff --git a/codex/chatgpt-capture-mvp/implementation/0001-chatgpt-capture-extension-scaffold.md b/codex/chatgpt-capture-mvp/implementation/0001-chatgpt-capture-extension-scaffold.md new file mode 100644 index 0000000..0b8a8c0 --- /dev/null +++ b/codex/chatgpt-capture-mvp/implementation/0001-chatgpt-capture-extension-scaffold.md @@ -0,0 +1,50 @@ +Create a minimal Manifest V3 browser extension for Chrome/Edge named "Project Thoth Capture". + +Goal: +Build only the extension skeleton for Sprint Task 1. Do not implement ChatGPT extraction yet. + +Requirements: +1. Use Manifest V3. +2. Add a browser toolbar button. +3. Add a background service worker. +4. When the toolbar button is clicked, log a message to the console: "Project Thoth capture clicked". +5. Include placeholder extension icons. +6. Include a README with local installation instructions for Chrome/Edge developer mode. +7. Keep the project simple and dependency-free. Use plain JavaScript, HTML only if needed, and no build tools. + +Project structure: + +project-thoth/applications/chatgpt-capture/ + manifest.json + src/ + background.js + icons/ + icon16.png + icon48.png + icon128.png + README.md + +Manifest details: +- name: "Project Thoth Capture" +- description: "Capture ChatGPT conversations for Project Thoth." +- version: "0.1.0" +- manifest_version: 3 +- permissions: ["activeTab", "scripting", "downloads"] +- host_permissions: + - "https://chatgpt.com/*" + - "https://chat.openai.com/*" +- action: + - default_title: "Save to Project Thoth" + - default_icon using the three icon files +- background: + - service_worker: "src/background.js" + +Implementation notes: +- The background service worker should register chrome.action.onClicked. +- On click, log the current tab URL and title if available. +- Do not inject scripts yet. +- Do not download files yet. +- Do not access page DOM yet. +- Add TODO comments for future capture flow. + +Return the complete file tree and the contents of each file. \ No newline at end of file diff --git a/codex/chatgpt-capture-mvp/implementation/0002-active-tab-capture.md b/codex/chatgpt-capture-mvp/implementation/0002-active-tab-capture.md new file mode 100644 index 0000000..b8018ff --- /dev/null +++ b/codex/chatgpt-capture-mvp/implementation/0002-active-tab-capture.md @@ -0,0 +1,50 @@ +Implement Task 2 — Active Tab Capture Flow for the Project Thoth ChatGPT Capture Connector MVP. + +Context: +- This is a Manifest V3 Chrome/Edge extension. +- Capture Connectors must only capture source material; they must not reason, call LLMs, generate metadata, or interact with the Project Thoth application yet. +- The MVP action is a browser toolbar button labeled “Save to Project Thoth.” +- The extension should use minimal permissions: activeTab, scripting, and downloads as needed later. +- This task focuses only on toolbar click → active tab lookup → script injection → service worker/content script message passing. + +Architecture requirements: +1. Use the Manifest V3 service worker as the orchestration point. +2. On toolbar button click, get the currently active tab. +3. Verify the current tab URL is ChatGPT: + - https://chatgpt.com/* + - https://chat.openai.com/* +4. Inject or execute a content script using chrome.scripting. +5. The content script should send a message back to the service worker with a placeholder capture payload. +6. The service worker should receive and log the payload. +7. Do not implement the real ChatGPT DOM extractor yet. +8. Do not implement Markdown normalization yet. +9. Do not implement downloads yet except leaving clear TODO boundaries. +10. Keep ChatGPT-specific extraction isolated for the next task. + +Expected placeholder payload: + +{ + sourcePlatform: "ChatGPT", + title: document.title || "ChatGPT Conversation", + url: window.location.href, + capturedAt: new Date().toISOString(), + messages: [] +} + +Files to update or create: +- manifest.json +- src/background.js +- src/chatgptExtractor.js if needed as a placeholder injected module +- README.md if local testing instructions need updating + +Acceptance criteria: +- Extension loads locally in Chrome/Edge. +- Toolbar button click triggers the service worker. +- On a ChatGPT tab, the content script executes successfully. +- A placeholder capture payload is returned to the service worker. +- On a non-ChatGPT tab, the extension does not inject and logs or displays a simple error. +- No LLM calls. +- No Project Thoth application dependency. +- No vault writes. + +Keep the implementation simple, readable, and modular so Task 3 can replace the placeholder with the real ChatGPT DOM extractor. \ No newline at end of file diff --git a/codex/chatgpt-capture-mvp/implementation/0003-dom-extractor.md b/codex/chatgpt-capture-mvp/implementation/0003-dom-extractor.md new file mode 100644 index 0000000..c77269f --- /dev/null +++ b/codex/chatgpt-capture-mvp/implementation/0003-dom-extractor.md @@ -0,0 +1,124 @@ +Implement Task 3 — ChatGPT DOM Extractor for the Project Thoth ChatGPT Capture Connector MVP. + +Context: +- This is a Manifest V3 Chrome/Edge extension. +- Task 1 extension skeleton is complete. +- Task 2 activeTab capture flow is complete. +- The service worker can inject or invoke a content script and receive a placeholder payload. +- This task replaces the placeholder capture logic with a ChatGPT-specific DOM extractor. +- Capture Connectors must only capture source material. Do not reason, summarize, classify, call LLMs, generate Project Thoth metadata, or write to the vault. + +Goal: +Create a ChatGPT-specific DOM extractor module that returns an ordered conversation capture payload with title, URL, capture timestamp, and messages. + +Files to update or create: +- src/chatgptExtractor.js +- src/background.js only if needed to integrate the extractor result +- README.md only if testing instructions need updating + +Expected output shape: + +{ + sourcePlatform: "ChatGPT", + title: "Detected conversation title", + url: "https://chatgpt.com/...", + capturedAt: "2026-07-08T...", + messages: [ + { + role: "user", + content: "..." + }, + { + role: "assistant", + content: "..." + } + ] +} + +Extractor requirements: +1. Extract the conversation title if available. + - Prefer a visible conversation title if one exists. + - Fall back to document.title. + - Fall back to "ChatGPT Conversation". + +2. Extract ordered message blocks. + - Preserve the page order of messages. + - Return messages in the same order they appear in the conversation. + - Ignore navigation, sidebar, composer/input box, buttons, menus, and unrelated UI chrome. + +3. Detect speaker role when possible. + - Detect user messages as role: "user". + - Detect assistant messages as role: "assistant". + - If role cannot be determined, use role: "unknown" rather than guessing too aggressively. + - Keep role detection logic isolated and easy to revise because ChatGPT DOM structure may change. + +4. Extract text content. + - Extract the readable content of each message. + - Trim leading/trailing whitespace. + - Collapse excessive blank lines where appropriate. + - Do not include copy buttons, feedback controls, model labels, timestamps, or hidden UI text. + +5. Preserve basic formatting where practical. + - Preserve paragraph breaks. + - Preserve markdown-like headings. + - Preserve bullet and numbered lists as readable text. + - Preserve code blocks using fenced markdown when detectable. + - Preserve inline code as readable text when detectable. + - Do not attempt perfect HTML-to-Markdown conversion in this task. + +6. Handle code blocks. + - Detect pre/code blocks inside assistant messages. + - Output fenced code blocks. + - Include a language tag if available from the DOM; otherwise use a plain triple-backtick fence. + - Avoid duplicating code block content in the surrounding extracted text. + +7. Be defensive against ChatGPT DOM instability. + - Use multiple selector strategies where reasonable. + - Prefer semantic attributes when available, such as data-testid, role, aria-label, or known message container patterns. + - Keep selectors centralized near the top of the extractor. + - Add comments explaining selector assumptions. + +8. Return useful failure states. + - If no messages are found, return a structured payload with messages: [] and an error field such as: + { + error: "No ChatGPT conversation messages were detected." + } + - Do not throw unhandled exceptions from normal extraction failure. + +9. No browser download logic in this task. +10. No markdown normalizer in this task. +11. No Project Thoth application dependency. +12. No LLM calls. + +Suggested module API: + +export function extractChatGPTConversation() { + return { + sourcePlatform: "ChatGPT", + title, + url: window.location.href, + capturedAt: new Date().toISOString(), + messages + }; +} + +Suggested helper functions: +- getConversationTitle() +- findMessageElements() +- detectMessageRole(element) +- extractMessageContent(element) +- extractNodeAsMarkdown(node) +- normalizeWhitespace(text) + +Integration: +- The activeTab capture flow should invoke extractChatGPTConversation() from the injected content context. +- The service worker should receive and log the returned payload. +- Existing Task 2 message passing should remain intact. + +Acceptance criteria: +- On an open ChatGPT conversation, clicking the extension button returns a payload with title, URL, capturedAt, and ordered messages. +- User and assistant roles are detected when possible. +- Message content is readable and excludes obvious UI chrome. +- Paragraphs, lists, headings, and code blocks remain usable in the extracted content. +- If no conversation is detected, the payload clearly reports that no messages were found. +- The implementation remains modular so Task 4 can consume the payload and convert it into canonical conversation.md format. \ No newline at end of file diff --git a/codex/chatgpt-capture-mvp/implementation/0003a-conversation-discovery.md b/codex/chatgpt-capture-mvp/implementation/0003a-conversation-discovery.md new file mode 100644 index 0000000..11fad7e --- /dev/null +++ b/codex/chatgpt-capture-mvp/implementation/0003a-conversation-discovery.md @@ -0,0 +1,34 @@ +Implement Task 3A — ChatGPT Conversation Discovery. + +Goal: +Refactor the extractor so it reliably discovers all conversation-bearing containers before attempting high-quality Markdown conversion. + +Primary requirement: +Completeness is more important than formatting. + +Deliverables: +- findConversationContainers() +- classifyContainer(container) +- extractContainerFallback(container) +- debug reporting + +Rules: +- Never silently skip a candidate container. +- If role cannot be determined, mark role: "unknown". +- If type cannot be determined, mark type: "unknown". +- Preserve rawText for every discovered container. +- Preserve a short htmlSnippet for unknown containers. +- Return debug counts: + - containersFound + - parsedMessages + - userMessages + - assistantMessages + - unknownContainers + +Do not solve tables, links, bold, lists, or Markdown formatting in this task. +Those belong to Task 3B. + +Acceptance criteria: +- Previously lost conversation sections are captured as either user, assistant, or unknown. +- Unknown sections are included in output rather than dropped. +- Debug output makes it clear what was discovered and what was not confidently parsed. \ No newline at end of file diff --git a/codex/chatgpt-capture-mvp/implementation/0003b-dom-to-markdown-conversion.md b/codex/chatgpt-capture-mvp/implementation/0003b-dom-to-markdown-conversion.md new file mode 100644 index 0000000..0a8c9d4 --- /dev/null +++ b/codex/chatgpt-capture-mvp/implementation/0003b-dom-to-markdown-conversion.md @@ -0,0 +1,87 @@ +Implement Task 3B — ChatGPT DOM-to-Markdown Conversion. + +Context: +- Task 3A discovers all conversation-bearing containers. +- 3A prioritizes completeness and returns containers/messages with role, type, rawText, htmlSnippet, and debug data. +- This task improves formatting quality for discovered ChatGPT message containers. +- Do not change the 3A discovery strategy unless absolutely necessary. + +Goal: +Convert discovered ChatGPT message DOM content into readable Markdown while preserving source structure. + +Primary rule: +Do not lose content. If structured conversion fails, fall back to rawText. + +Requirements: +1. Preserve paragraphs: + -

becomes one paragraph. + - Paragraphs separated by one blank line. + - Inline elements stay inline. + +2. Preserve inline formatting: + - / → **text** + - / → *text* + - inline → `text` + - → [text](url) + +3. Preserve lists: + -