Skip to content
Navigation menu
Search
Powered by Algolia
Search
Log in
Create account
DEV Community
Close
#
pytorch
Follow
Hide
Posts
Left menu
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
Right menu
Neural Networks: Weights, Activation, and Backpropagation
Aahan-Chauhan
Aahan-Chauhan
Aahan-Chauhan
Follow
Sep 6
Neural Networks: Weights, Activation, and Backpropagation
#
machinelearning
#
beginners
#
deeplearning
#
pytorch
Comments
Add Comment
5 min read
What Your Loss Function Actually Tells the Model: MSE, Cross-Entropy, and the softmax Bug That Trains Anyway
Wesam Khallaf — Author of PyTorch From Ground Up
Wesam Khallaf — Author of PyTorch From Ground Up
Wesam Khallaf — Author of PyTorch From Ground Up
Follow
Sep 8
What Your Loss Function Actually Tells the Model: MSE, Cross-Entropy, and the softmax Bug That Trains Anyway
#
pytorch
#
deeplearning
#
machinelearning
#
beginners
Comments
Add Comment
13 min read
Attention is simpler than you think - A hand-crafted superhero transformer
Tech-Aarvam
Tech-Aarvam
Tech-Aarvam
Follow
Sep 8
Attention is simpler than you think - A hand-crafted superhero transformer
#
ai
#
llm
#
pytorch
Comments
1
 comment
9 min read
What optimizer.step() Actually Does: One Step of SGD, Momentum and Adam by Hand
Wesam Khallaf — Author of PyTorch From Ground Up
Wesam Khallaf — Author of PyTorch From Ground Up
Wesam Khallaf — Author of PyTorch From Ground Up
Follow
Aug 28
What optimizer.step() Actually Does: One Step of SGD, Momentum and Adam by Hand
#
pytorch
#
deeplearning
#
machinelearning
#
beginners
Comments
1
 comment
10 min read
Backpropagation by Hand: Two Layers, a Pen, and Then Autograd Agrees
Wesam Khallaf — Author of PyTorch From Ground Up
Wesam Khallaf — Author of PyTorch From Ground Up
Wesam Khallaf — Author of PyTorch From Ground Up
Follow
Aug 23
Backpropagation by Hand: Two Layers, a Pen, and Then Autograd Agrees
#
pytorch
#
python
#
machinelearning
#
beginners
Comments
Add Comment
9 min read
Model DNA, Analyzed: Verifying 'From-Scratch' LLM Claims with Architecture, Tokenizer, and CKA (PyTorch)
ai maya
ai maya
ai maya
Follow
Aug 9
Model DNA, Analyzed: Verifying 'From-Scratch' LLM Claims with Architecture, Tokenizer, and CKA (PyTorch)
#
ai
#
llm
#
machinelearning
#
pytorch
Comments
Add Comment
6 min read
ResKV: Recovering Evicted Token Contributions via Residual KV Cache — LongBench 32/32
Chaeyeon Mia Lee
Chaeyeon Mia Lee
Chaeyeon Mia Lee
Follow
Aug 6
ResKV: Recovering Evicted Token Contributions via Residual KV Cache — LongBench 32/32
#
llm
#
machinelearning
#
performance
#
pytorch
Comments
Add Comment
6 min read
Your quantized model got worse, and nothing told you
AS
AS
AS
Follow
Aug 2
Your quantized model got worse, and nothing told you
#
flutter
#
dart
#
machinelearning
#
pytorch
3
 reactions
Comments
2
 comments
5 min read
Demystifying Shape Mismatches: How to Debug and Fix Tensor Dimensions in PyTorch
Oluwasegun Zyden
Oluwasegun Zyden
Oluwasegun Zyden
Follow
Jul 28
Demystifying Shape Mismatches: How to Debug and Fix Tensor Dimensions in PyTorch
#
python
#
machinelearning
#
pytorch
#
tutorial
Comments
Add Comment
4 min read
I Trained a 6.4M-Parameter Transformer From Scratch to Talk About Recipes
Medha
Medha
Medha
Follow
Jul 25
I Trained a 6.4M-Parameter Transformer From Scratch to Talk About Recipes
#
machinelearning
#
pytorch
#
llm
#
python
1
 reaction
Comments
Add Comment
5 min read
RoPE: How 2D Rotations Solved Transformer Long-Context
Masih Maafi
Masih Maafi
Masih Maafi
Follow
Jul 22
RoPE: How 2D Rotations Solved Transformer Long-Context
#
python
#
machinelearning
#
ai
#
pytorch
1
 reaction
Comments
Add Comment
4 min read
Running Qwen3 Through the ExecuTorch MLX Delegate: Up to 4.52x Faster on M1 Max
Nariaki Wada
Nariaki Wada
Nariaki Wada
Follow
Jul 22
Running Qwen3 Through the ExecuTorch MLX Delegate: Up to 4.52x Faster on M1 Max
#
python
#
pytorch
#
llm
#
applesilicon
Comments
Add Comment
7 min read
Testing PyTorch 2.13 MPS FlexAttention on M1 Max: Up to 7.83x Faster for Sparse Attention
Nariaki Wada
Nariaki Wada
Nariaki Wada
Follow
Jul 21
Testing PyTorch 2.13 MPS FlexAttention on M1 Max: Up to 7.83x Faster for Sparse Attention
#
python
#
pytorch
#
machinelearning
#
applesilicon
Comments
Add Comment
7 min read
GPU-Accelerating MSCRED with CUDA, im2col, GEMM, and a Custom PyTorch Extension
ranjithvutnoor
ranjithvutnoor
ranjithvutnoor
Follow
Aug 5
GPU-Accelerating MSCRED with CUDA, im2col, GEMM, and a Custom PyTorch Extension
#
pytorch
#
cuda
#
machinelearning
#
performance
4
 reactions
Comments
2
 comments
11 min read
Classifier-free guidance above 7.5 oversaturated our product renders
Elise Moreau
Elise Moreau
Elise Moreau
Follow
Jun 26
Classifier-free guidance above 7.5 oversaturated our product renders
#
machinelearning
#
computervision
#
pytorch
1
 reaction
Comments
Add Comment
4 min read
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
We're a place where coders share, stay up-to-date and grow their careers.
Log in
Create account