office docs
Ingest DOCX, PPTX, XLSX, and Notion/Confluence docs into structured Markdown for RAG pipelines.
Free
Works with the AI tools you already use
office docs
Example session with this skill installed
I have a large XLSX file with 5000 rows. I need to chunk it for a RAG system so that each chunk includes the column headers and keeps the sheet name in metadata. Show me the code.
- Read your context and instructions
- Compiled the office docs
I have implemented a table-aware chunking strategy for your XLSX file. This uses openpyxl for memory-efficient reading and pandas for Markdown conversion, ensuring headers are preserved in every block.
XLSX Table-Aware Chunker
import openpyxl
import pandas as pd
from dataclasses import dataclass
@dataclass
class SheetChunk:
sheet: str
range: str
markdown: str
row_start: int
row_end: int
def process_large_spreadsheet(path: str, rows_per_chunk: int = 50):
# read_only=True is critical for 5000+ row files
wb = openpyxl.load_workbook(path, data_only=True, read_only=True)
chunks = []
for sheet_name in wb.sheetnames:
df = pd.read_excel(path, sheet_name=sheet_name)
if df.empty:
continue
for start in range(0, len(df), rows_per_chunk):
end = min(start + rows_per_chunk, len(df))
slice_df = df.iloc[start:end]
# Convert to markdown to preserve structure for the LLM
md_content = slice_df.to_markdown(index=False)
chunks.append(SheetChunk(
sheet=sheet_name,
range=f"A{start+2}:Z{end+1}",
markdown=md_content,
row_start=start + 2,
row_end=end + 1
))
return chunks
# Example usage
chunks = process_large_spreadsheet("financial_data.xlsx")
print(f"Generated {len(chunks)} chunks with headers preserved.")
Next steps
- Install dependencies via
pip install openpyxl pandas tabulate. - Integrate the
SheetChunkobjects into your vector database upsert logic. - Use the
sheetandrangefields as metadata to allow the LLM to cite specific cells.
Connects securely to your tools. The creator never sees your data.
About this skill
The problem
Standard text loaders often flatten Office files into a mess of plain text, losing the headings, table structures, and slide notes that provide critical context for RAG systems. Large spreadsheets or slide decks can also overflow context windows if not chunked with structural awareness.
What it does
- Converts DOCX to Markdown while preserving heading hierarchies, lists, and tracked changes using mammoth and python-docx.
- Extracts PPTX content slide-by-slide, capturing both visual text and speaker notes while saving embedded images.
- Implements table-aware XLSX chunking with pandas and openpyxl, ensuring headers are repeated across chunks for data integrity.
- Automates structured exports from Notion, Confluence, and Quip into clean Markdown for vector ingestion.
- Identifies and extracts embedded images, cell comments, and document revisions as distinct metadata layers.
Frameworks & tools
Python, python-docx, python-pptx, openpyxl, pandas, mammoth, notion-client, atlassian-python-api, and markdownify.
Why this beats prompting it yourself
Extracting nested data like DOCX comments or XLSX formulas requires specific library knowledge and boilerplate that generic LLM prompts often hallucinate. This skill provides production-ready logic for handling large files in read-only modes and maintaining metadata paths that improve RAG retrieval accuracy.
Use cases
- Building a RAG pipeline for internal corporate knowledge bases.
- Automating technical documentation migration from Confluence to GitHub.
- Extracting structured data from complex financial spreadsheets for analysis.
- Syncing Notion workspace content into a vector database.
Known limitations
Requires specific Python libraries to be installed. Does not handle PDF files; use specialized extraction tools for those.
How to install
Works the same in every agent - Claude, Cursor, Codex, Copilot and 20+ more.
- 1
Download the ZIP
Free skills download straight away. Paid skills unlock right after purchase.
- 2
Unzip into your skills folder
Every agent reads skills from one folder on your machine. Drop the unzipped folder in there.
- 3
Ask your agent to use it
Restart the agent if it was already running. It picks the skill up automatically - no config needed.
Skills folder by agent
Click the path to copy it. Create the folder if it does not exist yet.
Reviews
No reviews yet
Be one of the first to try it. Every listed skill passes our trust checks below.
Security scanned
Passed our 8-point scan before listing
1 install
Downloaded by developers to date
Free forever
No account required to browse
Trust & safety
Security scanned
Verified clean 12 days ago
- Free to download with an account
Needs access to