Skip to content

Latest commit

 

History

25 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 

Repository files navigation

TensorTonic Solutions

Welcome to my TensorTonic solutions repository!

Here you'll find my solutions to various machine learning and deep learning problems from TensorTonic.

What is TensorTonic?

TensorTonic is a platform where you can implement core algorithms of Machine Learning from scratch.

This repository contains my personal solutions to these problems, automatically synchronized from the platform.

RAGHAV KACHROO's TensorTonic Solutions

Verified machine learning implementations completed on TensorTonic.

TensorTonic Verified Solutions

Problem Description Link
Dot Product Implement a multi-block CUDA dot-product reduction that combines partial sums into one scalar output. https://www.tensortonic.com/study-plans/cuda-basics/cuda/dot-product
GELU Implement exact GELU activation in CUDA with one thread per element and the device error-function intrinsic. https://www.tensortonic.com/study-plans/cuda-basics/cuda/gelu
Matrix-Vector Multiplication Implement row-major CUDA matrix-vector multiplication with one thread computing each output row. https://www.tensortonic.com/study-plans/cuda-basics/cuda/gemv
Hadamard Product Implement elementwise matrix multiplication in CUDA using a two-dimensional grid and row-major bounds-checked indexing. https://www.tensortonic.com/study-plans/cuda-basics/cuda/hadamard-product
Layer Normalization Implement fused row-wise LayerNorm in CUDA with shared-memory mean and variance reduction, affine scale, and bias. https://www.tensortonic.com/study-plans/cuda-basics/cuda/layer-norm
Leaky ReLU Implement Leaky ReLU activation in CUDA with one thread per element, bounds checks, and a configurable negative slope. https://www.tensortonic.com/study-plans/cuda-basics/cuda/leaky-relu
Matrix Addition Implement elementwise matrix addition in CUDA with a two-dimensional grid, row-major indexing, and bounds checks. https://www.tensortonic.com/study-plans/cuda-basics/cuda/matrix-addition
Matrix Multiplication Implement row-major matrix multiplication in CUDA with one thread per output element and inner-product accumulation. https://www.tensortonic.com/study-plans/cuda-basics/cuda/matrix-multiplication
Matrix Transpose Implement matrix transpose in CUDA with a two-dimensional launch grid, row-major buffers, and bounds-checked writes. https://www.tensortonic.com/study-plans/cuda-basics/cuda/matrix-transpose
Outer Product Compute a vector outer product in CUDA with a two-dimensional grid, row-major output, and bounds-checked indexing. https://www.tensortonic.com/study-plans/cuda-basics/cuda/outer-product
ReLU Implement ReLU activation in CUDA with one thread per element, bounds checks, and branch-efficient rectification. https://www.tensortonic.com/study-plans/cuda-basics/cuda/relu
RMS Normalization Implement row-wise RMS normalization in CUDA with sum-of-squares reduction, numerical stability, and learnable scaling. https://www.tensortonic.com/study-plans/cuda-basics/cuda/rms-norm
Sigmoid Implement sigmoid activation in CUDA with one thread per element, device exponential math, and bounds-checked memory access. https://www.tensortonic.com/study-plans/cuda-basics/cuda/sigmoid
Softmax Implement numerically stable CUDA softmax over a vector using global maximum and normalization reductions. https://www.tensortonic.com/study-plans/cuda-basics/cuda/softmax
Sum of Array Implement a multi-block CUDA sum reduction that combines partial block sums into one scalar output. https://www.tensortonic.com/study-plans/cuda-basics/cuda/sum-of-array
Swish Implement fused Swish or SiLU activation in CUDA with one thread per element and device exponential math. https://www.tensortonic.com/study-plans/cuda-basics/cuda/swish
Tanh Implement hyperbolic tangent activation in CUDA with one thread per element, device intrinsic math, and bounds checks. https://www.tensortonic.com/study-plans/cuda-basics/cuda/tanh
Vector Addition Implement bounds-checked pointwise vector addition in CUDA with one thread per output element. https://www.tensortonic.com/study-plans/cuda-basics/cuda/vector-addition
Vector Subtraction Implement bounds-checked pointwise vector subtraction in CUDA with one thread per output element. https://www.tensortonic.com/study-plans/cuda-basics/cuda/vector-subtract

View my verified ML profile: TensorTonic profile

About

My solutions to TensorTonic problems

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages