RagLeap
Guide · ragleap-tools

ragleap-tools package guide

Standalone, dependency-light tool implementations for LLM tool-calling. Exposes OpenAI/Gemini-style function-calling schemas.

pip install ragleap-tools

What this is (and isn't)

ragleap-tools provides Tool objects - a name, a description, a JSON Schema for parameters, and a safe handler function. It does not own a tool-calling execution loop (deciding when to call a tool, running it, feeding the result back to the model) - that's ragleap-agents' job, per the project roadmap's own split. Wire these tools into your own tool-calling code, or into ragleap-agents once that ships.

Quickstart

from ragleap_tools import STATELESS_TOOLS, CALCULATOR_TOOL
# Give these to your LLM provider's tools= parameter:
openai_tools = [t.to_openai_schema() for t in STATELESS_TOOLS]
gemini_tools = [t.to_gemini_schema() for t in STATELESS_TOOLS]
# When the model calls one, invoke the handler yourself:
result = CALCULATOR_TOOL.call(expression="2 + 2 * sqrt(16)")
print(result.success, result.result)  # True 10.0

The 12 stateless tools

No configuration needed - import and use directly.

  • CALCULATOR_TOOL - safe arithmetic (AST-based whitelist, never eval())
  • CURRENT_DATETIME_TOOL, ADD_TO_DATE_TOOL, DATE_DIFFERENCE_TOOL - date/time math
  • CONVERT_LENGTH_TOOL, CONVERT_WEIGHT_TOOL, CONVERT_TEMPERATURE_TOOL - unit conversion
  • PARSE_JSON_TOOL, PARSE_CSV_TOOL - structured data parsing
  • REGEX_EXTRACT_TOOL, WORD_COUNT_TOOL, TEXT_CASE_TRANSFORM_TOOL - text utilities

File operations (sandboxed, needs configuration)

from ragleap_tools import FileOpsConfig, make_file_tools
config = FileOpsConfig(root_dir="/path/to/a/safe/directory")
read_tool, write_tool, list_tool = make_file_tools(config)

Every operation is confined to root_dir - both ../ path traversal and symlink-based escapes are rejected (verified via real security tests, not just documented), not just naive string-prefix checking. There is no unsandboxed mode.

Document ingestion and search (optional, needs ragleap-rag)

pip install ragleap-tools[ingest]
from ragleap import RagLeap, ProviderConfig, EmbeddingConfig
from ragleap_tools import IngestConfig, make_ingest_tool, SearchConfig, make_search_tool
rag = RagLeap(database_url="...", primary=ProviderConfig(...), embedder=EmbeddingConfig(...))
ingest_tool = make_ingest_tool(IngestConfig(rag=rag))
search_tool = make_search_tool(SearchConfig(rag=rag))

ingest_document wraps ragleap-rag's already-tested ingest_text() - no new ingestion logic, just a tool schema on top of the real 28-format-capable pipeline. As of v0.1.1, it stores the filename as metadata ({"filename": filename}), enabling the per-document search below - v0.1.0 did not pass any metadata, which silently made per-document filtering impossible.

search_documents wraps ragleap-rag's already-tested retrieve() for hybrid vector+keyword search over previously ingested documents. Chunk dicts are returned unmodified - this tool doesn't assume ragleap-rag's exact field set. Pass an optional filename= to scope the search to a single document previously ingested via ingest_document:

result = search_tool.call(query="what was the Q3 revenue?", filename="q3-report.pdf")

ragleap-rag owns the actual ingestion and retrieval logic; these are thin adapters, same pattern ragleap-graph uses for its own optional ragleap-rag dependency.

Web search (BYOK, no extra dependencies)

from ragleap_tools import WebSearchConfig, TavilySearchProvider, make_web_search_tool
provider = TavilySearchProvider(api_key="...")  # or SerperSearchProvider(api_key="...")
search_tool = make_web_search_tool(WebSearchConfig(provider=provider))
result = search_tool.call(query="latest pgvector release", num_results=5)
# result.result == {"results": [{"title": ..., "url": ..., "snippet": ...}, ...], "count": 5}

Bring your own key: you construct the provider you want with your own api_key. There is no default provider and no environment-variable fallback. Uses only the standard library (urllib.request), so ragleap-tools still has zero required dependencies. SearchProvider is a small abstract class - implement search(query, num_results) to plug in any other search engine.

num_results is chosen by the model and costs your API quota, so it is clamped to 1-20.

Honest limitations: request shapes for both reference providers were checked against their current public docs, but neither has been called against a live account - treat them as best-effort until confirmed by someone with a real key. Results are text from arbitrary third-party pages, so treat them as untrusted input: this tool does not screen them for prompt injection. There is no caching, rate limiting or deduplication.

GitHub repository search (BYOK-optional, no extra dependencies)

from ragleap_tools import GitHubSearchConfig, make_github_search_tool
tool = make_github_search_tool(GitHubSearchConfig())  # token optional
# or: GitHubSearchConfig(token="...") for a higher rate limit
result = tool.call(query="language:python topic:llm", num_results=5, sort="stars")
# result.result == {"results": [{"full_name": ..., "url": ..., "description": ..., "stars": ..., "language": ...}, ...], "count": 5}

Works without a token, at GitHub's lower unauthenticated limit for the search resource (a live /rate_limit check on 2026-10-01 showed 10 for search, versus 60 for GitHub's core API; check /rate_limit for your own) - pass token= for a higher limit. No environment-variable fallback: if you want a token used, you pass it. Supports GitHub's real search qualifiers in the query string (language:, stars:, topic:, etc.), same as GitHub's own search UI. Scoped to repository search only, not code or issue search. Standard library only (urllib.request) - no new dependency.

Image description (BYOK, no extra dependencies)

from ragleap_tools import (
    FileOpsConfig, GeminiVisionProvider, VisionConfig, make_vision_tool,
)
provider = GeminiVisionProvider(api_key="...", model="...")  # you choose the model
tool = make_vision_tool(VisionConfig(
    provider=provider,
    sandbox=FileOpsConfig(root_dir="/path/to/images"),
))
result = tool.call(path="chart.png", prompt="What does this chart show?")
# result.result == {"description": "..."}

api_key and model are both required - no environment-variable fallback, no default provider, no default model. AnthropicVisionProvider works the same way (and also takes max_tokens, default 1024, which caps the description's length and cost). The tool only reads images inside the sandbox directory, using the same path-escape protection as the file tools (../ traversal, absolute paths and symlink escapes are rejected). The image type is detected from the file's bytes, not its extension. Defaults: 5,000,000 bytes per image (max_image_bytes) and 2,000 characters for the model-supplied prompt (max_prompt_chars). Gemini accepts JPEG, PNG and WebP here; Anthropic also accepts GIF. Standard library only (urllib.request) - no new dependency.

URLs are deliberately not accepted: letting a model name a URL to fetch carries the same SSRF risk as the HTTP-fetch tool listed below as out of scope. The returned description is text derived from an image that may be attacker-controlled - treat it as untrusted input (it is not screened for prompt injection).

GeminiVisionProvider was live-checked once on 2026-10-03 (one call on a small PNG, model gemini-3.6-flash): the request shape and response parsing worked. JPEG and WebP input and Gemini's error responses were not live-checked. AnthropicVisionProvider has not been called against a live account - its request shape was checked against current public documentation only, so treat it as best-effort until confirmed live.

Deliberately out of scope

Each of these needs its own security-focused design pass, not a rushed inclusion here:

  • Code execution - a real sandboxing/resource-limit design decision, not something to bolt on alongside a calculator.
  • HTTP fetch - letting an LLM request arbitrary URLs carries real SSRF risk, same care level as code execution.
  • Database/business-system connectors (SQL, CRM, payment processors, etc.) - some of what this ecosystem already has elsewhere (e.g. a live payment processor) would be a materially different risk if exposed to LLM tool-calling without deliberate guardrails (dry-run modes, confirmation steps, scoped permissions).

Status

v0.4.0. 133 tests, all passing, including real security verification for the two risk-sensitive tools (calculator's code-injection rejection, file ops' path-traversal and symlink-escape rejection) - not just documented as safe, actually tested against real attack vectors.

Not live-verified: the web search providers (Tavily, Serper), the Anthropic vision provider, and authenticated GitHub requests - request shapes were checked against current public documentation only. The Gemini vision provider was live-checked once (see above).

License

MIT

This page mirrors README.md on GitHub. GitHub is the source of truth and may be newer.

All ragleap-tools guides Package page All guides