updated docs strings and added README.md

2026-03-08 17:59:56 +05:30
parent 0fbf0ca0f0
commit de7d04eb1a
26 changed files with 546 additions and 406 deletions
--- a/mcp_docs/modules/omniread.pdf.json
+++ b/mcp_docs/modules/omniread.pdf.json
@@ -2,14 +2,14 @@
  "module": "omniread.pdf",
  "content": {
    "path": "omniread.pdf",
-    "docstring": "PDF format implementation for OmniRead.\n\n---\n\n## Summary\n\nThis package provides **PDF-specific implementations** of the core OmniRead\ncontracts defined in `omniread.core`.\n\nUnlike HTML, PDF handling requires an explicit client layer for document\naccess. This package therefore includes:\n- PDF clients for acquiring raw PDF data\n- PDF scrapers that coordinate client access\n- PDF parsers that extract structured content from PDF binaries\n\nPublic exports from this package represent the supported PDF pipeline\nand are safe for consumers to import directly when working with PDFs.\n\n---\n\n## Public API\n\n    FileSystemPDFClient\n    PDFScraper\n    PDFParser\n\n---",
+    "docstring": "# Summary\n\nPDF format implementation for OmniRead.\n\nThis package provides **PDF-specific implementations** of the core OmniRead\ncontracts defined in `omniread.core`.\n\nUnlike HTML, PDF handling requires an explicit client layer for document\naccess. This package therefore includes:\n\n- PDF clients for acquiring raw PDF data.\n- PDF scrapers that coordinate client access.\n- PDF parsers that extract structured content from PDF binaries.\n\nPublic exports from this package represent the supported PDF pipeline\nand are safe for consumers to import directly when working with PDFs.\n\n---\n\n# Public API\n\n- `FileSystemPDFClient`\n- `PDFScraper`\n- `PDFParser`\n\n---",
    "objects": {
      "FileSystemPDFClient": {
        "name": "FileSystemPDFClient",
        "kind": "class",
        "path": "omniread.pdf.FileSystemPDFClient",
        "signature": "<bound method Alias.signature of Alias('FileSystemPDFClient', 'omniread.pdf.client.FileSystemPDFClient')>",
-        "docstring": "PDF client that reads from the local filesystem.\n\nNotes:\n    **Guarantees:**\n\n        - This client reads PDF files directly from the disk and returns their raw binary contents",
+        "docstring": "PDF client that reads from the local filesystem.\n\nNotes:\n    **Guarantees:**\n\n        - This client reads PDF files directly from the disk and returns\n          their raw binary contents.",
        "members": {
          "fetch": {
            "name": "fetch",
@@ -25,7 +25,7 @@
        "kind": "class",
        "path": "omniread.pdf.PDFScraper",
        "signature": "<bound method Alias.signature of Alias('PDFScraper', 'omniread.pdf.scraper.PDFScraper')>",
-        "docstring": "Scraper for PDF sources.\n\nNotes:\n    **Responsibilities:**\n\n        - Delegates byte retrieval to a PDF client and normalizes output into Content\n        - Preserves caller-provided metadata\n\n    **Constraints:**\n    \n        - The scraper: Does not perform parsing or interpretation, does not assume a specific storage backend",
+        "docstring": "Scraper for PDF sources.\n\nNotes:\n    **Responsibilities:**\n\n        - Delegates byte retrieval to a PDF client and normalizes output\n          into `Content`.\n        - Preserves caller-provided metadata.\n\n    **Constraints:**\n\n        - The scraper does not perform parsing or interpretation.\n        - Does not assume a specific storage backend.",
        "members": {
          "fetch": {
            "name": "fetch",
@@ -41,7 +41,7 @@
        "kind": "class",
        "path": "omniread.pdf.PDFParser",
        "signature": "<bound method Alias.signature of Alias('PDFParser', 'omniread.pdf.parser.PDFParser')>",
-        "docstring": "Base PDF parser.\n\nNotes:\n    **Responsibilities:**\n\n        - This class enforces PDF content-type compatibility and provides the extension point for implementing concrete PDF parsing strategies\n\n    **Constraints:**\n\n        - Concrete implementations must: Define the output type `T`, implement the `parse()` method",
+        "docstring": "Base PDF parser.\n\nNotes:\n    **Responsibilities:**\n\n        - This class enforces PDF content-type compatibility and provides\n          the extension point for implementing concrete PDF parsing strategies.\n\n    **Constraints:**\n\n        - Concrete implementations must define the output type `T` and\n          implement the `parse()` method.",
        "members": {
          "supported_types": {
            "name": "supported_types",
@@ -55,7 +55,7 @@
            "kind": "function",
            "path": "omniread.pdf.PDFParser.parse",
            "signature": "<bound method Alias.signature of Alias('parse', 'omniread.pdf.parser.PDFParser.parse')>",
-            "docstring": "Parse PDF content into a structured output.\n\nReturns:\n    T:\n        Parsed representation of type `T`.\n\nRaises:\n    Exception:\n        Parsing-specific errors as defined by the implementation.\n\nNotes:\n    **Responsibilities:**\n\n        - Implementations must fully interpret the PDF binary payload and return a deterministic, structured output"
+            "docstring": "Parse PDF content into a structured output.\n\nReturns:\n    T:\n        Parsed representation of type `T`.\n\nRaises:\n    Exception:\n        Parsing-specific errors as defined by the implementation.\n\nNotes:\n    **Responsibilities:**\n\n        - Implementations must fully interpret the PDF binary payload and\n          return a deterministic, structured output."
          }
        }
      },
@@ -64,7 +64,7 @@
        "kind": "module",
        "path": "omniread.pdf.client",
        "signature": null,
-        "docstring": "PDF client abstractions for OmniRead.\n\n---\n\n## Summary\n\nThis module defines the **client layer** responsible for retrieving raw PDF\nbytes from a concrete backing store.\n\nClients provide low-level access to PDF binaries and are intentionally\ndecoupled from scraping and parsing logic. They do not perform validation,\ninterpretation, or content extraction.\n\nTypical backing stores include:\n- Local filesystems\n- Object storage (S3, GCS, etc.)\n- Network file systems",
+        "docstring": "# Summary\n\nPDF client abstractions for OmniRead.\n\nThis module defines the **client layer** responsible for retrieving raw PDF\nbytes from a concrete backing store.\n\nClients provide low-level access to PDF binaries and are intentionally\ndecoupled from scraping and parsing logic. They do not perform validation,\ninterpretation, or content extraction.\n\nTypical backing stores include:\n\n- Local filesystems\n- Object storage (S3, GCS, etc.)\n- Network file systems",
        "members": {
          "Any": {
            "name": "Any",
@@ -98,14 +98,14 @@
            "name": "BasePDFClient",
            "kind": "class",
            "path": "omniread.pdf.client.BasePDFClient",
-            "signature": "<bound method Class.signature of Class('BasePDFClient', 26, 54)>",
-            "docstring": "Abstract client responsible for retrieving PDF bytes\nfrom a specific backing store (filesystem, S3, FTP, etc.).\n\nNotes:\n    **Responsibilities:**\n\n        - Implementations must accept a source identifier appropriate to the backing store, return the full PDF binary payload, and raise retrieval-specific errors on failure",
+            "signature": "<bound method Class.signature of Class('BasePDFClient', 25, 57)>",
+            "docstring": "Abstract client responsible for retrieving PDF bytes.\n\nRetrieves bytes from a specific backing store (filesystem, S3, FTP, etc.).\n\nNotes:\n    **Responsibilities:**\n\n        - Implementations must accept a source identifier appropriate to\n          the backing store.\n        - Return the full PDF binary payload.\n        - Raise retrieval-specific errors on failure.",
            "members": {
              "fetch": {
                "name": "fetch",
                "kind": "function",
                "path": "omniread.pdf.client.BasePDFClient.fetch",
-                "signature": "<bound method Function.signature of Function('fetch', 37, 54)>",
+                "signature": "<bound method Function.signature of Function('fetch', 40, 57)>",
                "docstring": "Fetch raw PDF bytes from the given source.\n\nArgs:\n    source (Any):\n        Identifier of the PDF location, such as a file path, object storage key, or remote reference.\n\nReturns:\n    bytes:\n        Raw PDF bytes.\n\nRaises:\n    Exception:\n        Retrieval-specific errors defined by the implementation."
              }
            }
@@ -114,14 +114,14 @@
            "name": "FileSystemPDFClient",
            "kind": "class",
            "path": "omniread.pdf.client.FileSystemPDFClient",
-            "signature": "<bound method Class.signature of Class('FileSystemPDFClient', 57, 92)>",
-            "docstring": "PDF client that reads from the local filesystem.\n\nNotes:\n    **Guarantees:**\n\n        - This client reads PDF files directly from the disk and returns their raw binary contents",
+            "signature": "<bound method Class.signature of Class('FileSystemPDFClient', 60, 96)>",
+            "docstring": "PDF client that reads from the local filesystem.\n\nNotes:\n    **Guarantees:**\n\n        - This client reads PDF files directly from the disk and returns\n          their raw binary contents.",
            "members": {
              "fetch": {
                "name": "fetch",
                "kind": "function",
                "path": "omniread.pdf.client.FileSystemPDFClient.fetch",
-                "signature": "<bound method Function.signature of Function('fetch', 67, 92)>",
+                "signature": "<bound method Function.signature of Function('fetch', 71, 96)>",
                "docstring": "Read a PDF file from the local filesystem.\n\nArgs:\n    path (Path):\n        Filesystem path to the PDF file.\n\nReturns:\n    bytes:\n        Raw PDF bytes.\n\nRaises:\n    FileNotFoundError:\n        If the path does not exist.\n    ValueError:\n        If the path exists but is not a file."
              }
            }
@@ -133,7 +133,7 @@
        "kind": "module",
        "path": "omniread.pdf.parser",
        "signature": null,
-        "docstring": "PDF parser base implementations for OmniRead.\n\n---\n\n## Summary\n\nThis module defines the **PDF-specific parser contract**, extending the\nformat-agnostic `BaseParser` with constraints appropriate for PDF content.\n\nPDF parsers are responsible for interpreting binary PDF data and producing\nstructured representations suitable for downstream consumption.",
+        "docstring": "# Summary\n\nPDF parser base implementations for OmniRead.\n\nThis module defines the **PDF-specific parser contract**, extending the\nformat-agnostic `BaseParser` with constraints appropriate for PDF content.\n\nPDF parsers are responsible for interpreting binary PDF data and producing\nstructured representations suitable for downstream consumption.",
        "members": {
          "Generic": {
            "name": "Generic",
@@ -161,7 +161,7 @@
            "kind": "class",
            "path": "omniread.pdf.parser.ContentType",
            "signature": "<bound method Alias.signature of Alias('ContentType', 'omniread.core.content.ContentType')>",
-            "docstring": "Supported MIME types for extracted content.\n\nNotes:\n    **Guarantees:**\n\n        - This enum represents the declared or inferred media type of the content source\n        - It is primarily used for routing content to the appropriate parser or downstream consumer",
+            "docstring": "Supported MIME types for extracted content.\n\nNotes:\n    **Guarantees:**\n\n        - This enum represents the declared or inferred media type of the\n          content source.\n        - It is primarily used for routing content to the appropriate\n          parser or downstream consumer.",
            "members": {
              "HTML": {
                "name": "HTML",
@@ -198,7 +198,7 @@
            "kind": "class",
            "path": "omniread.pdf.parser.BaseParser",
            "signature": "<bound method Alias.signature of Alias('BaseParser', 'omniread.core.parser.BaseParser')>",
-            "docstring": "Base interface for all parsers.\n\nNotes:\n    **Guarantees:**\n\n        - A parser is a self-contained object that owns the Content it is responsible for interpreting\n        - Consumers may rely on early validation of content compatibility and type-stable return values from `parse()`\n\n    **Responsibilities:**\n\n        - Implementations must declare supported content types via `supported_types`\n        - Implementations must raise parsing-specific exceptions from `parse()`\n        - Implementations must remain deterministic for a given input",
+            "docstring": "Base interface for all parsers.\n\nNotes:\n    **Guarantees:**\n\n        - A parser is a self-contained object that owns the `Content` it is\n          responsible for interpreting.\n        - Consumers may rely on early validation of content compatibility\n          and type-stable return values from `parse()`.\n\n    **Responsibilities:**\n\n        - Implementations must declare supported content types via `supported_types`.\n        - Implementations must raise parsing-specific exceptions from `parse()`.\n        - Implementations must remain deterministic for a given input.",
            "members": {
              "supported_types": {
                "name": "supported_types",
@@ -219,7 +219,7 @@
                "kind": "function",
                "path": "omniread.pdf.parser.BaseParser.parse",
                "signature": "<bound method Alias.signature of Alias('parse', 'omniread.core.parser.BaseParser.parse')>",
-                "docstring": "Parse the owned content into structured output.\n\nReturns:\n    T:\n        Parsed, structured representation.\n\nRaises:\n    Exception:\n        Parsing-specific errors as defined by the implementation.\n\nNotes:\n    **Responsibilities:**\n\n        - Implementations must fully consume the provided content and return a deterministic, structured output"
+                "docstring": "Parse the owned content into structured output.\n\nReturns:\n    T:\n        Parsed, structured representation.\n\nRaises:\n    Exception:\n        Parsing-specific errors as defined by the implementation.\n\nNotes:\n    **Responsibilities:**\n\n        - Implementations must fully consume the provided content and\n          return a deterministic, structured output."
              },
              "supports": {
                "name": "supports",
@@ -241,8 +241,8 @@
            "name": "PDFParser",
            "kind": "class",
            "path": "omniread.pdf.parser.PDFParser",
-            "signature": "<bound method Class.signature of Class('PDFParser', 24, 61)>",
-            "docstring": "Base PDF parser.\n\nNotes:\n    **Responsibilities:**\n\n        - This class enforces PDF content-type compatibility and provides the extension point for implementing concrete PDF parsing strategies\n\n    **Constraints:**\n\n        - Concrete implementations must: Define the output type `T`, implement the `parse()` method",
+            "signature": "<bound method Class.signature of Class('PDFParser', 22, 62)>",
+            "docstring": "Base PDF parser.\n\nNotes:\n    **Responsibilities:**\n\n        - This class enforces PDF content-type compatibility and provides\n          the extension point for implementing concrete PDF parsing strategies.\n\n    **Constraints:**\n\n        - Concrete implementations must define the output type `T` and\n          implement the `parse()` method.",
            "members": {
              "supported_types": {
                "name": "supported_types",
@@ -255,8 +255,8 @@
                "name": "parse",
                "kind": "function",
                "path": "omniread.pdf.parser.PDFParser.parse",
-                "signature": "<bound method Function.signature of Function('parse', 43, 61)>",
-                "docstring": "Parse PDF content into a structured output.\n\nReturns:\n    T:\n        Parsed representation of type `T`.\n\nRaises:\n    Exception:\n        Parsing-specific errors as defined by the implementation.\n\nNotes:\n    **Responsibilities:**\n\n        - Implementations must fully interpret the PDF binary payload and return a deterministic, structured output"
+                "signature": "<bound method Function.signature of Function('parse', 43, 62)>",
+                "docstring": "Parse PDF content into a structured output.\n\nReturns:\n    T:\n        Parsed representation of type `T`.\n\nRaises:\n    Exception:\n        Parsing-specific errors as defined by the implementation.\n\nNotes:\n    **Responsibilities:**\n\n        - Implementations must fully interpret the PDF binary payload and\n          return a deterministic, structured output."
              }
            }
          }
@@ -267,7 +267,7 @@
        "kind": "module",
        "path": "omniread.pdf.scraper",
        "signature": null,
-        "docstring": "PDF scraping implementation for OmniRead.\n\n---\n\n## Summary\n\nThis module provides a PDF-specific scraper that coordinates PDF byte\nretrieval via a client and normalizes the result into a `Content` object.\n\nThe scraper implements the core `BaseScraper` contract while delegating\nall storage and access concerns to a `BasePDFClient` implementation.",
+        "docstring": "# Summary\n\nPDF scraping implementation for OmniRead.\n\nThis module provides a PDF-specific scraper that coordinates PDF byte\nretrieval via a client and normalizes the result into a `Content` object.\n\nThe scraper implements the core `BaseScraper` contract while delegating\nall storage and access concerns to a `BasePDFClient` implementation.",
        "members": {
          "Any": {
            "name": "Any",
@@ -295,7 +295,7 @@
            "kind": "class",
            "path": "omniread.pdf.scraper.Content",
            "signature": "<bound method Alias.signature of Alias('Content', 'omniread.core.content.Content')>",
-            "docstring": "Normalized representation of extracted content.\n\nNotes:\n    **Responsibilities:**\n\n        - A `Content` instance represents a raw content payload along with minimal contextual metadata describing its origin and type\n        - This class is the primary exchange format between Scrapers, Parsers, and Downstream consumers",
+            "docstring": "Normalized representation of extracted content.\n\nNotes:\n    **Responsibilities:**\n\n        - A `Content` instance represents a raw content payload along with\n          minimal contextual metadata describing its origin and type.\n        - This class is the primary exchange format between scrapers,\n          parsers, and downstream consumers.",
            "members": {
              "raw": {
                "name": "raw",
@@ -332,7 +332,7 @@
            "kind": "class",
            "path": "omniread.pdf.scraper.ContentType",
            "signature": "<bound method Alias.signature of Alias('ContentType', 'omniread.core.content.ContentType')>",
-            "docstring": "Supported MIME types for extracted content.\n\nNotes:\n    **Guarantees:**\n\n        - This enum represents the declared or inferred media type of the content source\n        - It is primarily used for routing content to the appropriate parser or downstream consumer",
+            "docstring": "Supported MIME types for extracted content.\n\nNotes:\n    **Guarantees:**\n\n        - This enum represents the declared or inferred media type of the\n          content source.\n        - It is primarily used for routing content to the appropriate\n          parser or downstream consumer.",
            "members": {
              "HTML": {
                "name": "HTML",
@@ -369,14 +369,14 @@
            "kind": "class",
            "path": "omniread.pdf.scraper.BaseScraper",
            "signature": "<bound method Alias.signature of Alias('BaseScraper', 'omniread.core.scraper.BaseScraper')>",
-            "docstring": "Base interface for all scrapers.\n\nNotes:\n    **Responsibilities:**\n\n        - A scraper is responsible ONLY for fetching raw content (bytes) from a source. It must not interpret or parse it\n        - A scraper is a stateless acquisition component that retrieves raw content from a source and returns it as a `Content` object\n        - Scrapers define how content is obtained, not what the content means\n        - Implementations may vary in transport mechanism, authentication strategy, retry and backoff behavior\n\n    **Constraints:**\n\n        - Implementations must not parse content, modify content semantics, or couple scraping logic to a specific parser",
+            "docstring": "Base interface for all scrapers.\n\nNotes:\n    **Responsibilities:**\n\n        - A scraper is responsible ONLY for fetching raw content (bytes)\n          from a source. It must not interpret or parse it.\n        - A scraper is a stateless acquisition component that retrieves raw\n          content from a source and returns it as a `Content` object.\n        - Scrapers define how content is obtained, not what the content means.\n        - Implementations may vary in transport mechanism, authentication\n          strategy, retry and backoff behavior.\n\n    **Constraints:**\n\n        - Implementations must not parse content, modify content semantics,\n          or couple scraping logic to a specific parser.",
            "members": {
              "fetch": {
                "name": "fetch",
                "kind": "function",
                "path": "omniread.pdf.scraper.BaseScraper.fetch",
                "signature": "<bound method Alias.signature of Alias('fetch', 'omniread.core.scraper.BaseScraper.fetch')>",
-                "docstring": "Fetch raw content from the given source.\n\nArgs:\n    source (str):\n        Location identifier (URL, file path, S3 URI, etc.)\n    metadata (Optional[Mapping[str, Any]], optional):\n        Optional hints for the scraper (headers, auth, etc.)\n\nReturns:\n    Content:\n        Content object containing raw bytes and metadata.\n\nRaises:\n    Exception:\n        Retrieval-specific errors as defined by the implementation.\n\nNotes:\n    **Responsibilities:**\n\n        - Implementations must retrieve the content referenced by `source` and return it as raw bytes wrapped in a `Content` object"
+                "docstring": "Fetch raw content from the given source.\n\nArgs:\n    source (str):\n        Location identifier (URL, file path, S3 URI, etc.).\n\n    metadata (Optional[Mapping[str, Any]], optional):\n        Optional hints for the scraper (headers, auth, etc.).\n\nReturns:\n    Content:\n        Content object containing raw bytes and metadata.\n\nRaises:\n    Exception:\n        Retrieval-specific errors as defined by the implementation.\n\nNotes:\n    **Responsibilities:**\n\n        - Implementations must retrieve the content referenced by `source`\n          and return it as raw bytes wrapped in a `Content` object."
              }
            }
          },
@@ -385,7 +385,7 @@
            "kind": "class",
            "path": "omniread.pdf.scraper.BasePDFClient",
            "signature": "<bound method Alias.signature of Alias('BasePDFClient', 'omniread.pdf.client.BasePDFClient')>",
-            "docstring": "Abstract client responsible for retrieving PDF bytes\nfrom a specific backing store (filesystem, S3, FTP, etc.).\n\nNotes:\n    **Responsibilities:**\n\n        - Implementations must accept a source identifier appropriate to the backing store, return the full PDF binary payload, and raise retrieval-specific errors on failure",
+            "docstring": "Abstract client responsible for retrieving PDF bytes.\n\nRetrieves bytes from a specific backing store (filesystem, S3, FTP, etc.).\n\nNotes:\n    **Responsibilities:**\n\n        - Implementations must accept a source identifier appropriate to\n          the backing store.\n        - Return the full PDF binary payload.\n        - Raise retrieval-specific errors on failure.",
            "members": {
              "fetch": {
                "name": "fetch",
@@ -400,8 +400,8 @@
            "name": "PDFScraper",
            "kind": "class",
            "path": "omniread.pdf.scraper.PDFScraper",
-            "signature": "<bound method Class.signature of Class('PDFScraper', 22, 77)>",
-            "docstring": "Scraper for PDF sources.\n\nNotes:\n    **Responsibilities:**\n\n        - Delegates byte retrieval to a PDF client and normalizes output into Content\n        - Preserves caller-provided metadata\n\n    **Constraints:**\n    \n        - The scraper: Does not perform parsing or interpretation, does not assume a specific storage backend",
+            "signature": "<bound method Class.signature of Class('PDFScraper', 20, 77)>",
+            "docstring": "Scraper for PDF sources.\n\nNotes:\n    **Responsibilities:**\n\n        - Delegates byte retrieval to a PDF client and normalizes output\n          into `Content`.\n        - Preserves caller-provided metadata.\n\n    **Constraints:**\n\n        - The scraper does not perform parsing or interpretation.\n        - Does not assume a specific storage backend.",
            "members": {
              "fetch": {
                "name": "fetch",