office docs

    1

    Ingest DOCX, PPTX, XLSX, and Notion/Confluence docs into structured Markdown for RAG pipelines.

    Free

    1 installSecurity scanned

    Works with the AI tools you already use

    Claude CodeClaude CodeCursorCursorCodex CLICodex CLIMuseMuseDotsDots+22 more

    office docs

    Example session with this skill installed

    I have a large XLSX file with 5000 rows. I need to chunk it for a RAG system so that each chunk includes the column headers and keeps the sheet name in metadata. Show me the code.

    • Read your context and instructions
    • Compiled the office docs

    I have implemented a table-aware chunking strategy for your XLSX file. This uses openpyxl for memory-efficient reading and pandas for Markdown conversion, ensuring headers are preserved in every block.

    XLSX Table-Aware Chunker

    import openpyxl
    import pandas as pd
    from dataclasses import dataclass
    
    @dataclass
    class SheetChunk:
        sheet: str
        range: str
        markdown: str
        row_start: int
        row_end: int
    
    def process_large_spreadsheet(path: str, rows_per_chunk: int = 50):
        # read_only=True is critical for 5000+ row files
        wb = openpyxl.load_workbook(path, data_only=True, read_only=True)
        chunks = []
    
        for sheet_name in wb.sheetnames:
            df = pd.read_excel(path, sheet_name=sheet_name)
            if df.empty:
                continue
    
            for start in range(0, len(df), rows_per_chunk):
                end = min(start + rows_per_chunk, len(df))
                slice_df = df.iloc[start:end]
                
                # Convert to markdown to preserve structure for the LLM
                md_content = slice_df.to_markdown(index=False)
                
                chunks.append(SheetChunk(
                    sheet=sheet_name,
                    range=f"A{start+2}:Z{end+1}",
                    markdown=md_content,
                    row_start=start + 2,
                    row_end=end + 1
                ))
        return chunks
    
    # Example usage
    chunks = process_large_spreadsheet("financial_data.xlsx")
    print(f"Generated {len(chunks)} chunks with headers preserved.")
    

    Next steps

    • Install dependencies via pip install openpyxl pandas tabulate.
    • Integrate the SheetChunk objects into your vector database upsert logic.
    • Use the sheet and range fields as metadata to allow the LLM to cite specific cells.

    Connects securely to your tools. The creator never sees your data.

    About this skill

    The problem

    Standard text loaders often flatten Office files into a mess of plain text, losing the headings, table structures, and slide notes that provide critical context for RAG systems. Large spreadsheets or slide decks can also overflow context windows if not chunked with structural awareness.

    What it does

    • Converts DOCX to Markdown while preserving heading hierarchies, lists, and tracked changes using mammoth and python-docx.
    • Extracts PPTX content slide-by-slide, capturing both visual text and speaker notes while saving embedded images.
    • Implements table-aware XLSX chunking with pandas and openpyxl, ensuring headers are repeated across chunks for data integrity.
    • Automates structured exports from Notion, Confluence, and Quip into clean Markdown for vector ingestion.
    • Identifies and extracts embedded images, cell comments, and document revisions as distinct metadata layers.

    Frameworks & tools

    Python, python-docx, python-pptx, openpyxl, pandas, mammoth, notion-client, atlassian-python-api, and markdownify.

    Why this beats prompting it yourself

    Extracting nested data like DOCX comments or XLSX formulas requires specific library knowledge and boilerplate that generic LLM prompts often hallucinate. This skill provides production-ready logic for handling large files in read-only modes and maintaining metadata paths that improve RAG retrieval accuracy.

    Use cases

    • Building a RAG pipeline for internal corporate knowledge bases.
    • Automating technical documentation migration from Confluence to GitHub.
    • Extracting structured data from complex financial spreadsheets for analysis.
    • Syncing Notion workspace content into a vector database.

    Known limitations

    Requires specific Python libraries to be installed. Does not handle PDF files; use specialized extraction tools for those.

    How to install

    Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.

    ~30 seconds
    1. 1

      Download the ZIP

      Free skills download straight away. Paid skills unlock right after purchase.

    2. 2

      Unzip into your skills folder

      Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.

    3. 3

      Ask your agent to use it

      Restart the agent if it was already running. It picks the skill up automatically - no config needed.

    Skills folder by agent

    Click the path to copy it. Create the folder if it does not exist yet.

    Reviews

    No reviews yet

    Be one of the first to try it. Every listed skill passes our trust checks below.

    Security scanned

    Passed our 8-point scan before listing

    1 install

    Downloaded by developers to date

    Free forever

    No account required to browse

    Trust & safety

    Security scanned

    Verified clean 12 days ago

    • Free to download with an account

    Needs access to

    Openxmlformats
    Atlassian
    Github

    Listed12 days ago

    What's inside

    Frequently Asked Questions