Problem Statement: Every month, my daughter's shool Parent-Teacher Council (PTC), just like every other school in USA, meets to discuss important issues related to kids, their leraning, teachers, the school environment, budgetting, etc. As a secretary to the board, my job is to keep the minutes of the meeting, and get it approved in the next meeting. You would think keeping meeting minutes is such a trivial job now a days with Zoom, Google-Meet, etc, supporting AI powered meeting minutes generation, but the School & the School Districts are not so fast in embracing these newer technologies due to governance, legacy subscriptions, and budgetting constraints. In our case, the School District has a legacy Zoom subscription and a policy that does not allow automatic meeting minutes.
Solution: Build an Agentic AI sytem that, given a meeting recording, generates meeting minutes, creates a document of preferred type, and stores or uploads it at a preferred location.
Why Agentic AI? While at first glance this might seem like a glorified automation, and in its current form that's somewhat true, the potential to make this more flexible by leverging the skills of the LLMs exists: (i) the encoding format of the input audio file need not be fixed, let the LLM figure out how to deal with any format; (ii) the file format for the MoMs need not be fixed to google doc, let the LLM figure out how to generate if other formats are asked for (potentially LLM generating the code on the fly with appropriate prompt changes TBD); (iii) the upload destination can be anything (School PTC Website, Box, AWS S3, ...); let LLM figure out what tools to use, and potentiall come up with the required code as required (TBD). Hence an Agentic AI system with differnt agents for each task is deliberately chosen for these future enhancements.
Why Input is a JSON File: The longer term goal is to expose an API that that takes the required arguments and does the job, so this tool can be easily integrated with the school's PTC Website, but this is a future enhancement.
An agentic AI pipeline that takes a Zoom recording (.mp4 / .m4a), transcribes it, generates structured meeting minutes, and publishes them to Google Drive — fully automated, end-to-end.
Built with Python · CrewAI · OpenAI Whisper · Google Docs & Drive APIs.
Runs in Jupyter Notebook or Google Colab.
┌─────────────────────────────────────────────────────────────────────────────┐
│ PTC MoMs — Agent Pipeline │
└─────────────────────────────────────────────────────────────────────────────┘
input.json
┌──────────────────────┐
│ media_file_path │ .mp4 or .m4a Zoom recording
│ meeting_date │ YYYY-MM-DD string
│ drive_folder_id │ Google Drive destination folder ID
│ credentials_path │ OAuth2 credentials JSON path
└──────────┬───────────┘
│
▼
╔══════════════════════╗
║ 1. Orchestrator ║ • Loads and parses input.json
║ Agent ║ • Validates every field (file existence,
╚══════════╦═══════════╝ extension, date format, credentials)
║ validated config
▼
╔══════════════════════╗
║ 2. Transcript ║ • Extracts audio from .mp4 if needed
║ Agent ║ • Runs OpenAI Whisper locally (CPU/GPU)
╚══════════╦═══════════╝ • Returns full verbatim transcript
║ verbatim transcript text
▼
╔══════════════════════╗
║ 3. Secretary ║ • Reads transcript with an LLM
║ Agent ║ • Produces six structured sections:
╚══════════╦═══════════╝ Executive Summary (≤50 words)
║ Motions Passed
║ Action Items (person → task)
║ Decisions Taken
║ Open Issues
║ Key Messages
║ markdown meeting minutes
▼
╔══════════════════════╗
║ 4. Google Docs ║ • Creates a new Google Document
║ Creator Agent ║ • Applies HEADING_1 / HEADING_2 styles
╚══════════╦═══════════╝ • Returns doc ID and shareable URL
║ doc_id + doc_url
▼
╔══════════════════════╗
║ 5. Drive Uploader ║ • Moves the doc into the specified
║ Agent ║ Google Drive folder
╚══════════╦═══════════╝ • Returns confirmation + shareable link
║
▼
✅ Meeting Minutes live in Google Drive
Each stage is also a standalone notebook (notebooks/0N_*.ipynb) that can be run and debugged independently.
PTC-MoMs-CrewAI/
│
├── main_pipeline.ipynb ← ▶ Run this for the full end-to-end pipeline
│
├── notebooks/
│ ├── 01_orchestrator.ipynb ← Stage 1 – input validation
│ ├── 02_transcript_agent.ipynb ← Stage 2 – Whisper transcription
│ ├── 03_secretary_agent.ipynb ← Stage 3 – meeting minutes (LLM)
│ ├── 04_docs_creator_agent.ipynb ← Stage 4 – Google Docs creation
│ └── 05_drive_uploader_agent.ipynb← Stage 5 – Google Drive upload
│
├── src/
│ ├── config.py ← Pydantic input validation model
│ ├── llm_config.py ← LLM factory (swap providers in .env)
│ ├── state.py ← JSON-backed state shared across notebooks
│ └── tools/
│ ├── transcription_tool.py ← CrewAI tool wrapping Whisper
│ ├── gdocs_tool.py ← CrewAI tool wrapping Google Docs API
│ └── gdrive_tool.py ← CrewAI tool wrapping Google Drive API
│
├── requirements.txt
├── .env.example ← Copy → .env and fill in your keys
├── input_template.json ← Copy → input.json and fill in your paths
└── pipeline_state.json ← Auto-created at runtime (inter-notebook state)
| Requirement | Notes |
|---|---|
| Python 3.10+ | 3.11 recommended |
ffmpeg |
Required for .mp4 audio extraction. Install via your OS package manager. |
| OpenAI API key | Or any supported LLM provider (see Switching LLM) |
| Google Cloud project | With Google Docs API and Google Drive API enabled |
| OAuth2 credentials JSON | Downloaded from Google Cloud Console |
# macOS
brew install ffmpeg
# Ubuntu / Debian
sudo apt install ffmpeg
# Windows — download from https://ffmpeg.org/download.html- Go to Google Cloud Console → APIs & Services → Library
- Enable Google Docs API
- Enable Google Drive API
- Go to APIs & Services → Credentials → Create Credentials → OAuth client ID
- Choose Desktop app, download the JSON → this is your
credentials.json
Why a dedicated venv?
crewaisupports Python 3.10–3.13 only. If JupyterLab is installed via Homebrew
(or any other method that bundles its own Python), its kernel may be Python 3.14+
and will fail to installcrewai. The steps below create a Python 3.11 venv and
register it as a named Jupyter kernel so JupyterLab uses the right interpreter.
# 1. Clone the repo (or copy the folder)
git clone <repo-url>
cd PTC-MoMs-CrewAI
# 2. Create a venv with Python 3.11 (binary name on macOS Homebrew is python3.11)
python3.11 -m venv venv
source venv/bin/activate # Windows: venv\Scripts\activate
# 3. Install all dependencies (including ipykernel)
pip install -r requirements.txt
# 4. Register this venv as a Jupyter kernel
python -m ipykernel install --user \
--name ptc-moms-py311 \
--display-name "PTC MoMs (Python 3.11)"
# 5. Launch JupyterLab (can be the system/Homebrew one — kernel is separate)
jupyter labOnce JupyterLab opens:
- Open
main_pipeline.ipynb - In the top-right corner click the kernel name → Change Kernel
- Select "PTC MoMs (Python 3.11)"
- Do the same for any individual notebook you open
The first code cell in every notebook verifies the Python version at runtime and
raises a clear error if the wrong kernel is selected.
Google Colab: Skip steps 1–5. The first cell in every notebook runs
pip install -r requirements.txtautomatically, and Colab's Python (3.10) is compatible.
cp .env.example .envOpen .env and set your LLM API key:
# .env
LLM_PROVIDER=openai
LLM_MODEL=gpt-4o
LLM_TEMPERATURE=0.1
OPENAI_API_KEY=sk-... # ← required for default OpenAI setup
# Whisper settings (see Transcription section below)
WHISPER_MODEL_SIZE=base
WHISPER_DEVICE=cpucp input_template.json input.jsonEdit input.json with your meeting details:
| Parameter | Type | Description |
|---|---|---|
media_file_path |
string | Absolute path to the Zoom recording. Supported: .mp4, .m4a |
meeting_date |
string | Date the meeting was held. Format: YYYY-MM-DD |
google_drive_folder_id |
string | The folder ID from the Google Drive URL: drive.google.com/drive/folders/<ID> |
oauth2_credentials_path |
string | Absolute path to the OAuth2 credentials.json downloaded from Google Cloud |
Finding your Drive folder ID:
Open the folder in Google Drive. The URL looks like
https://drive.google.com/drive/folders/1BxiMVs0XRA5nFMdKvBdBZjgmUUqptlbs
The long alphanumeric string at the end is the folder ID.
Open main_pipeline.ipynb in Jupyter Lab / Jupyter Notebook / Google Colab and run all cells top-to-bottom.
# Local Jupyter
jupyter lab main_pipeline.ipynbThe notebook will:
- Install dependencies
- Load and validate
input.json - Transcribe the recording (may take several minutes for long files)
- Generate meeting minutes
- Create and upload the Google Doc
- Print the final shareable URL
First run — Google auth: A browser window will open asking you to grant
access to Google Docs and Drive. After consent, atoken.jsonis saved next
to yourcredentials.jsonso subsequent runs skip this step.
In Colab, authentication is handled automatically viagoogle.colab.auth.
Each notebooks/0N_*.ipynb notebook can run on its own. Notebooks exchange data via pipeline_state.json (created automatically).
Run in order: Or jump to any stage by filling in the
OVERRIDE_* variables at the top of the notebook.
01_orchestrator.ipynb ──┐
02_transcript_agent.ipynb ├─► pipeline_state.json ──► each notebook reads/writes here
03_secretary_agent.ipynb │
04_docs_creator_agent.ipynb
05_drive_uploader_agent.ipynb
Standalone example — run only the Secretary agent on an existing transcript:
Open 03_secretary_agent.ipynb and set:
# Near the top of the notebook
OVERRIDE_TRANSCRIPT = """
Good morning everyone. Let us call the meeting to order...
[paste your transcript here]
"""Then run all cells.
Change two lines in .env — no code changes required.
# Anthropic Claude
LLM_PROVIDER=anthropic
LLM_MODEL=claude-3-5-sonnet-20241022
ANTHROPIC_API_KEY=sk-ant-...
# Groq (fast inference)
LLM_PROVIDER=groq
LLM_MODEL=llama3-70b-8192
GROQ_API_KEY=gsk_...
# Ollama (local, no API key needed)
LLM_PROVIDER=ollama
LLM_MODEL=llama3Supported providers are defined in src/llm_config.py. Adding a new one is a two-line change in that file.
The transcription agent uses faster-whisper running locally — no audio leaves your machine.
It uses CTranslate2 (not PyTorch/numba), so there is no LLVM dependency and it installs cleanly on all platforms.
WHISPER_MODEL_SIZE |
Speed | Accuracy | RAM (CPU) |
|---|---|---|---|
tiny |
fastest | lowest | ~400 MB |
base |
fast | good | ~500 MB |
small |
medium | better | ~1 GB |
medium |
slow | great | ~2.5 GB |
large-v2 |
slowest | best | ~5 GB |
large-v3 |
slowest | best+ | ~5 GB |
# .env — CPU (default, no GPU required)
WHISPER_MODEL_SIZE=base
WHISPER_DEVICE=cpu
# WHISPER_COMPUTE_TYPE=int8 ← auto-set for CPU; override if needed
# .env — GPU (NVIDIA CUDA)
WHISPER_MODEL_SIZE=medium
WHISPER_DEVICE=cuda
# WHISPER_COMPUTE_TYPE=float16 ← auto-set for CUDA; override if neededFor
.mp4files, audio is extracted automatically to a temporary.wavfile usingmoviepy+ffmpeg, then deleted after transcription.
The Secretary agent always produces six sections in this order:
| Section | Description |
|---|---|
| Executive Summary | ≤ 50 words — high-level overview of the meeting |
| Motions Passed | Formally moved and seconded resolutions |
| Action Items | [Person / Team]: [what they must do] |
| Decisions Taken | Agreements reached during the meeting |
| Open Issues | Topics raised but left unresolved or deferred |
| Key Messages | Important statements and announcements from speakers |
| Symptom | Fix |
|---|---|
crewai>=0.80.0 — No matching distribution found |
Wrong kernel (Python 3.14+). Switch to "PTC MoMs (Python 3.11)" — see Installation above |
Python X.Y is not supported (cell 1 error) |
Same — select the correct kernel in JupyterLab before running |
| Kernel "PTC MoMs (Python 3.11)" not listed | Run python -m ipykernel install --user --name ptc-moms-py311 --display-name "PTC MoMs (Python 3.11)" inside the activated venv |
FileNotFoundError: credentials.json |
Check the oauth2_credentials_path in input.json |
Google Docs API has not been used… |
Enable the API in Google Cloud Console |
Token has expired |
Delete token.json (next to credentials.json) and re-run |
| Whisper runs very slowly | Use a smaller model (tiny or base) or set WHISPER_DEVICE=cuda |
moviepy import error |
Run pip install moviepy or ensure ffmpeg is on your PATH |
| Empty transcript | File may be silent or corrupt — verify playback before running |
MIT
{ "media_file_path": "/absolute/path/to/zoom_recording.mp4", "meeting_date": "2024-04-15", "google_drive_folder_id": "1BxiMVs0XRA5nFMdKvBdBZjgmUUqptlbs", "oauth2_credentials_path":"/absolute/path/to/credentials.json" }