LogoDocumentation
AI as a Service (AIaaS)

AIaaS Introduction

AI for Science. Built for Researchers. Run securely in e-INFRA CZ.

CERIT‑SC, a core component of e‑INFRA CZ, operates an on‑premise AI platform that provides researchers with secure, high‑performance, and interoperable AI tools. This platform runs on cutting‑edge NVIDIA DGX‑H100/B200/B300‑class systems, delivering robust computational power for high‑speed inference. The environment hosts open large language and generative models, accessible via the Open WebUI interface or standard OpenAI‑compatible APIs.

New: Matrix Community ChannelMatrix is now available for LLM service users to share knowledge and best practices, exchange experiences, and stay up to date with service status.

Key Features (Inference)

  • Secure, on‑premise LLM & generative‑AI platform – Our models run securely on the e‑INFRA CZ infrastructure. Queries and responses are not logged by external providers, ensuring that your research and sensitive data remain within our environment. With the exception of internet searching, nothing leaves our local infrastructure.
  • Supports privacy‑sensitive research – compliance with institutional and legal requirements. This makes our services ideal for handling sensitive data.

For more details, see AI Data privacy section

What the Platform Provides

CategoryHighlights
ComputeNVIDIA DGX‑H100/B200/B300. Petaflop‑class GPU performance.
Key Models

Our portfolio includes advanced models tailored for programming, image generation, code generation, tool use, and agentic workflows. Featured models include Kimi, GLM, DeepSeek, Qwen,… For enhanced security, these models operate entirely offline without internet access View other available models

Model status page

API token personal usage

Models’ MetadataYou can retrieve models’ metadata through the /model/info endpoint.
AccessMUNI students and employees and MetaCentrum users

Key AI Services (Inference)

ServiceLinkDescription
Chat (Open‑WebUI)WebUI chat

A full‑featured conversational interface (similar to ChatGPT) offering advanced features and explicit language model selection.

Text Work: Translations, summaries, analysis, and generation of program code.

Multimodality: Image generation (including editing) and content recognition in images (e.g., extracting a serial number from a photo).

Tools: Searching the internet, GitHub, arXiv, and a Python sandbox for running code in the browser and data analytics.

RAG (Knowledge): Searching within a document attached to the chat works very well.

OpenAI-compatible APIOpenAI APIUse the OpenAI‑compatible endpoint to integrate AI into your scripts, pipelines, or services.
MCP serversMCP serversWe provide several MCP servers that are available through the https://llm.ai.e-infra.cz API. A particular MCP server is available at https://llm.ai.e-infra.cz/[servername]/mcp.
AI Coding AssistantsAI Coding Assistants IntegrationBy connecting your tools to our backend, you can leverage high‑performance models for coding tasks (ClaudeCode, OpenCode, VSCode,… support.
DeepSiteDeepSite (vibe‑coding)A generative tool that creates webpages and applications (HTML/CSS/JS) based on a simple text description. Excellent for design proposals, mockups, or quick web concepts.
AI in Jupyter NotebooksJupyterHub integrationIntegration of an AI Assistant directly into the Jupyter Lab environment Notebook Intelligence). Used for fixing code (R or Python), generating new snippets, and conversational assistance within your coding projects.
n8n platformn8n Agentsn8n is an open‑source, low‑code workflow automation platform with AI Agent functionality
DeepSec vulnerability scannerDeepSecAI-powered vulnerability scanner designed to run in your own infrastructure. It performs on-demand security reviews of entire codebases — including large-scale repositories — by combining fast regex-based candidate detection with deep AI investigation.
Documentation ChatBotdocs.e-infra.cz and other documentation sitesA Retrieval‑Augmented Generation (RAG) system implemented across e‑INFRA CZ documentation. Answers specific questions and acts as a problem solver based on our knowledge base (e.g., “How to run MATLAB?”).
Matrix Discussion ChannelMatrixReal-time chat for LLM service users to ask questions, share knowledge, and stay updated.

Reference

Read more details on our e‑INFRA Blog at https://blog.e-infra.cz/

Picture used from https://blogs.nvidia.com/blog/difference-deep-learning-training-inference-ai/

publicity banner

On this page

einfra banner