≡ menu
× Menu

cnnmmd

overview


This is a collection of sample tools for connecting various applications.

The samples include tools for chat, tools for autonomous characters, tools for agents, and tools for network construction.

Since the modules are packaged into a library, please use it as one of your reference resources when generating your own code.



*
An explanation of this video (including rights information) is provided below:
・
dscmov

subject


This sample collection is for the following people:

・
Those who generate their own code [※1][※2]

*1
This sample is neither for developers (who can understand the code) nor for users (who only use the tool).
*2
Nowadays, using text generation models allows you to handle everything yourself, from creating the code to configuring the machine. However, without a set of guidelines, the more you continue generating text, the more inconsistent the styles become, making maintenance difficult. But if the generator clearly states the guidelines they followed, it might become possible to have a dialogue with other generators at that level—this is one attempt at resource sharing in the generation age.

resource


Here are some examples of the following toolset:

・
A suite of tools for chat [※1]
・
Tools for autonomous characters [※2]
・
A suite of tools for agents [※3]
・
A suite of tools for network construction [※4]

*1
It connects various characters (virtual (2D images to 3D models, web browser/desktop/VR/MR) to real (dolls/figures)) with various AI engines (speech recognition/speech synthesis/text generation/sentiment analysis/object recognition, local/remote/cloud).
*2
We provide environments for reinforcement learning (2D to 3D, classical optimization/bipedal locomotion) and evolutionary computation (optimization of coordination between parts).
*3
We provide execution of various programming languages, conversion of various documents, searching of various databases, creation and searching of vector data, and automated execution of web browsers.
*4
We provide various internet servers (name server, mail server (sending/receiving), web server, and proxy server) and their client suites. We also offer an X Window environment for GUI clients.

The modules are organized into libraries as follows:

・
Separation of operating environments, centralization of communications [※5]
・
Easy customization and plugin creation [※6]
・
Flexible configuration to suit different usage scenarios [※7]

*5
By introducing containers, we prevent environment conflicts. Communication between containers will be handled using a server-client architecture with the introduction of an intermediary server.
*6
Depending on the level of customization, we will implement static libraries, dynamic libraries, or dynamic code generation. The independence of each plugin is also guaranteed through namespaces.
*7
We will allow users to choose between a GUI-based approach that visualizes dependencies and a CLI-based approach that allows for iterative processing, depending on the situation.

image
image
image
image
image
image
image
image
image
image

environment


At this time, we have verified the connectivity between the following character sets and AI engine sets:

○
Character set (= client-side device + application)
・
Workflow (ComfyUI)

Mobile device + OS (iOS / Android) / PC + OS (Windows / macOS / Linux)
・
web browser

Mobile devices + OS (iOS / Android) / PCs + OS (Windows / macOS / Linux)
・
Desktop application (Electron)

+ PC + OS (Windows / macOS / Linux)
・
Game engine (Unity)

+ PC + OS (Windows / macOS)
・
VR/MR applications (Unity / Virt-A-Mate)

+ VR/MR equipment (Quest) + PC (Windows)
・
Microcontroller: Bare metal (M5Stack)

+ LCD/Microphone (M5Stack)
・
Figure (Nendoroid Doll)

+ Microcontroller: Bare metal (M5Stack) + Microphone/Camera/Speaker/Motor (M5Stack)
○
AI engine suite (i.e., AI-related model operation/service utilization)
・
Speech recognition model:

- Local (Transformers: wav2vec2-large-japanese / OpenAI: Whisper / Moonshine Voice)

• Service (OpenAI: Whisper)
・
Speech synthesis model:

・Local (VOICEVOX / Style-Bert-VITS2 / MioTTS-Inference / Irodori-TTS)

• Service ( NIJI Voice [API] ) [※1]
・
Text generation model:

・Local (llama.cpp: (gpt-oss / Gemma 3n / Gemma 4) / Transformers: Vecteus-v1)

• Services (OpenAI: GPT / NovelAI)
・
Sentiment analysis model:

- Local (Transformers: luke-japanese-large-sentiment-analysis-wrime)

• Services (OpenAI: GPT)
・
Image generation model:

- Local (Diffusers: Stable Diffusion / ComfyUI: Stable Diffusion, Anima, ...) [※2]

• Services (OpenAI: DALL-E / NovelAI)
・
Video generation model:

- Local (ComfyUI: Wan 2.2, MiniMax H3, ...) [※2]
・
Object recognition model:

- Local (OpenCV: Haar Cascades)
・
Environment for reinforcement learning/evolutionary computation:

• Local (Gymnasium / Evolution Gym / MuJoCo / Meta Motivo)
・
State machines: For language models:

• Local (LangGraph)
・
Vector store creation/search:

- Local (SentenceTransformer)
・
Runtime environment: For language models:

• Local (llama-server / Ollama)
◯
Relay server (a server that relays communication between servers / between a server and a client)
・
PC + OS (Windows (WLS2 (Linux (Ubuntu))))

+ Containers (Docker / Docker Desktop)
・
PC (Mac) + OS (macOS)

+ Containers (Docker Desktop)
・
PC + OS (Linux (Ubuntu))

+ Containers (Docker)
・
Microcontroller (Raspberry Pi) + OS (Linux (Ubuntu))

+ Containers (Docker)
○
Execution infrastructure (= programs/documents/databases/internet infrastructure)
・
Program execution sandbox:

- Local (RISC-V / Bash / Haskell / Prolog / Ruby / Scheme / Python + PyTorch / TypeScript)
・
Document conversion/database search sandbox:

・Local (pandoc (Markdown, HTML, TeX, PDF, DOCX, ...) / tidy (HTML) / xsltproc (XML/XSLT) / jq (JSON) / python:rdflib (RDF/SPARQL) / sqlite3 (RDB/SQL))
・
Internet server cluster:

- Local (Bind / Postfix / Dovecot / Apache / Nginx)
・
Internet client group:

- Local (CLI / GUI (Thunderbird / Firefox))
・
X Window System:

- Local (X Window server & RDP client / X Window clients)
・
Web browser automated execution environment:

• Local (Playwright)

*
The character group (the devices and applications on which the characters operate) is verifying connectivity with media ranging from virtual to real, including images, videos, 3D, VR, MR, and figurines. The AI ​​engine group (AI-related model operation/service utilization) is testing speech recognition, speech synthesis, text generation, sentiment analysis, image generation, and object recognition, supporting CPU/GPU operation and API calls. The relay server connecting them operates in environments ranging from PCs to microcontrollers.

*
The client/server/engine code is provided as a sample only (sufficient functionality is not guaranteed) -- this tool provides the creation of a relay server handler group and its workflow (i.e., a flow that connects the handler group as nodes).
*1
The Niji Voice API usage handler has been removed from the repository because the rights to the source audio (speech) are unclear (2025.11.08):
・
https://x.com/JAU_Official/status/1986664719441404392
*2
ComfyUI allows you to generate images and videos using any model by downloading a set of models and applying a workflow (such as the official sample) (and adding any custom nodes if needed).

Restrictions/Caution


At this time, the following limitations/notes apply:

◯
constraints
・
A container is required to use the tool [※1].
・
Currently, the connection for conversational communication is unstable [※2].
・
Backward compatibility is not guaranteed.
◯
Note
・
There is a linked article for adult-oriented generative models/client applications [※3].

*1
To allow the use of various AI engines without restrictions and to prevent conflicts between applications, we use OS container technology (Docker).
*2
We've made it easy to implement client code even on low-performance devices (such as bare-metal microcontrollers), so currently, communication is limited to HTTP polling. Streaming will also be implemented in a simplified version first (WebSocket/WebRTC will be added later). If the timing is off, the connection between the server and client will fail. In such cases, you will need to stop and restart the relay server or the GUI workflow creation environment (ComfyUI).
*3
Services aimed at the general public have a public nature, so it's inevitable that various restrictions will be imposed. While this ensures safety, it also narrows creative expression—one of the advantages of building a local environment is the absence of such restrictions. Here, we present generative models/client applications for adults without distinction, as long as the technology has appropriate characteristics—articles containing links to them are clearly indicated, and minors should look for the following mark : !!! NSFW / R18 !!!

right


The proprietary code for this tool is distributed under the following license:

・
MIT License

The resources (models, libraries, and applications) used by this tool are owned by their respective owners. Resources requiring special handling are clearly indicated in the description of each plugin. [※1]


*1
Some of the models used by this tool's engine (such as speech synthesis models, image generation models, and video generation models) can utilize voices, images, and videos trained by individuals. This tool is not intended for use beyond personal use, and is not intended for any actions that infringe on portrait rights, copyrights, or moral rights of authors.