Easy bidirectional Python/C/C++/JavaScript interop + compute kernel generation for CUDA, OpenCL, Vulkan, WebGPU, CPU SIMD

Tattletale is an AI inference engine written in Nim.
Nim compiles to C, C++ or Javascript.
It can call Python (and be called from Python) seamlessly.
Nim has powerful metaprogramming capabilities and offers direct AST manipulation through Nim itself.
It is possible to implement compilers, code generator and even CuTe layout algebra or an autodiff in Nim compile-time macros.
Tattletale uses them to generate high-performance GPU kernels at compile time for CUDA, OpenCL, Vulkan, and WebGPU, all from the same source.


What if C, Python, and Lisp had a child?

Nearly every major data tool in the Python ecosystem is built the same way: a fast C, C++, Rust, or Fortran core wrapped with Python bindings. NumPy, SciPy, Pandas, Polars, PyTorch, xgboost, the pattern repeats. What if you could combine the convenience of Python, with a powerful type system, and the speed of C?

(And maybe without waiting 20 minutes to a couple of hours for Cutlass and Flash Attention to compile)

I am building Tattletale, an AI inference engine, with all the scientific computing primitives needed to reach state-of-the-art performance in terms of plain speed, concurrent serving, multi-GPU support, and embeddability. Single binary, no dependency hell.

The foundation is Nim. Python-like syntax, compiles to C, C++, or JavaScript. Types. Zero-overhead bidirectional FFI with C and C++. A Python bridge that works both ways and can call Python from Nim or Nim from Python. Direct AST manipulation at compile time for metaprogrammingin the Lisp tradition.

This combination means: one codebase generates GPU kernels for CUDA, OpenCL, Vulkan, and WebGPU at compile time via Nim macros.
No multi-stage compilation, ninja or nvcc needed where you deploy your code.
The same source produces a Python-callable library, a C-embeddable binary, and even JavaScript or WASM targets.

This talk covers:

  • What Nim looks like in practice
  • How the bidirectional Python bridge works — calling Python from Nim and Nim from Python
  • Showcasing Nim's Domain Specific Languages: einsum as macros, neural network DSLs
  • Compile-time kernel generation: a Nim macro that emits CUDA + Vulkan from the same Nim source
  • CuTe Layout Algebra in Nim
  • How to seamlessly integrate with LibTorch C+_+ and even catch C++ exceptions.

No Nim knowledge required. Familiarity with FFI or GPU programming helps but is not essential.

Mamy

Mamy is a cryptography engineer, deep learning engineer and has designed too many threadpools to count. He spends his time shipping any math-heavy research to production, optimizing it along the way to better use the actual hardware they run on.