04667-A-More-Elegant-Word-Vector-Model-I-simpler-glove.pdf 04669-A-More-Unique-Word-Vector-Model-II-Modeling-Language.pdf 04671-A-More-Unique-Word-Vector-Model-III-Models-Describing-Correlation.pdf 04675-A-More-Unique-Word-Vector-Model-IV-Solving-the-Model.pdf 04677-A-More-Unique-Word-Vector-Model-V-Interesting-Results.pdf 04681-A-More-Unique-Word-Vector-Model-Part-6-Code-Sharing-and-Conclusion.pdf 04695-CRF-In-A-Nutshell.pdf 04718-The-Method-of-Characteristics-for-First-Order-Partial-Differential-Equations.pdf 04733-From-Hard-Truncation-and-Softening-of-Loss-to-Focal-Loss.pdf 04765-A-Brief-Reading-of-Attention-is-All-You-Need-Introduction-Code.pdf 04823-Sharing-a-Slide-Fancy-Natural-Language-Processing.pdf 05067-Sharing-10-Million-Level-Baidu-Zhidao-Corpus.pdf 05112-Another-New-Year-s-Feast-From-K-Means-to-Capsule.pdf 05155-Three-Flavors-of-Capsule-Matrix-Capsule-and-EM-Routing.pdf 05239-From-Maximum-Likelihood-to-EM-Algorithm-A-Consistent-Way-of-Understanding.pdf 05253-Variational-Autoencoders-Part-1-So-That-s-How-It-Is.pdf 05332-Poetry-Robot-Based-on-CNN-and-VAE-Random-Poetry-Generation.pdf 05343-Variational-Autoencoders-Part-2-From-a-Bayesian-Perspective.pdf 05383-Variational-Autoencoders-III-Why-Does-This-Work.pdf 05409-A-CNN-based-Reading-Comprehension-Question-Answering-Model-DGCNN.pdf 05448-The-Minimum-Entropy-Principle-I-The-Principle-of-Unsupervised-Learning.pdf 05476-The-Principle-of-Minimum-Entropy-II-Lexicon-Construction-via-Decisive-Cutting.pdf 05505-Spectral-Classification-Model-Based-on-Conv1D-One-Dimensional-Sequence-Classification.pdf 05525-Efficient-Implementation-of-the-Apriori-Algorithm-with-Numpy.pdf 05542-A-Concise-Introduction-to-Conditional-Random-Fields-CRF-with-a-Pure-Keras-Implementation.pdf 05577-The-Principle-of-Minimum-Entropy-III-Flying-Elephant-Across-the-River-Sentence-Templates-and-Language-Structure.pdf 05597-An-NLP-Library-Based-on-the-Principle-of-Minimum-Entropy-nlp-zero.pdf 05607-Simple-Python-Implementation-of-Gillespie-Simulation.pdf 05617-Notes-on-Noise-Contrastive-Estimation-The-Beauty-of-a-Winding-Path.pdf 05643-RNNs-and-ODEs-Seemingly-Different-but-Spiritually-Allied-An-Introduction-to-Fancy-RNNs.pdf 05655-Optimization-Algorithms-from-a-Dynamics-Perspective-I-From-SGD-to-Momentum-Acceleration.pdf 05693-From-SamplePairing-to-mixup-The-Magical-Regularization-Term.pdf 05716-A-Unified-Understanding-of-Generative-Models-via-Variational-Inference-VAE-GAN-AAE-ALI.pdf 05743-Sentence-Similarity-Model-Based-on-GRU-and-AM-Softmax.pdf 05765-Making-Keras-Cooler-Ingenious-Layers-and-Fancy-Callbacks.pdf 05776-NICE-Basic-Concepts-and-Implementation-of-Flow-Models.pdf 05807-Flowing-with-Flow-RealNVP-and-GlowThe-Inheritance-and-Sublimation-of-Flow-Models.pdf 05861-Playing-with-Keras-Automatic-Title-Generation-with-seq2seq.pdf 05879-Making-Keras-Even-Cooler-Niche-Custom-Optimizers.pdf 05887-Variational-Autoencoders-IV-A-One-Step-Clustering-Scheme.pdf 05977-f-VAEs-The-Marriage-of-Glow-and-VAEs.pdf 06016-Introduction-to-f-GAN-A-Production-Workshop-for-GAN-Models.pdf 06024-Mutual-Information-in-Deep-Learning-Unsupervised-Feature-Extraction.pdf 06051-Lipschitz-Constraints-in-Deep-Learning-Generalization-and-Generative-Models.pdf 06088-Variational-Autoencoder-Minimizing-Prior-Distribution-Maximizing-Mutual-Information.pdf 06096-Rethinking-the-Determinant-of-Non-Square-Matrices.pdf 06110-RSGAN-The-Turing-Test-Concept-in-Adversarial-Models.pdf 06131-In-Memory-of-Jin-Yong-May-You-Soar-on-Asteroid-10930.pdf 06139-WGAN-div-An-Obscure-WGAN-Pit-Filler.pdf 06163-A-GAN-without-Lipschitz-Constraints-and-Gradient-Vanishing-Interested.pdf 06181-From-Variational-Encoding-and-Information-Bottleneck-to-Normal-Distribution-On-the-Importance-of-Forgetting.pdf 06191-The-Minimum-Entropy-Principle-IV-Birds-of-a-Feather-Flock-Together-From-Libraries-to-Word-Vectors.pdf 06214-BiGAN-QP-A-Simple-and-Clear-Encoding-Generative-Model.pdf 06234-Optimization-Algorithms-from-a-Dynamics-Perspective-II-Adaptive-Learning-Rate-Algorithms.pdf 06240-Learning-List-Important-Recent-Advances-in-GAN-Papers.pdf 06261-Optimization-Algorithms-from-a-Dynamics-Perspective-III-A-More-Holistic-View.pdf 06270-A-Couplet-Robot-Based-on-CNN-and-Sequence-Tagging.pdf 06280-From-Wasserstein-Distance-and-Duality-Theory-to-WGAN.pdf 06311-Make-Keras-Cooler-Arbitrary-Outputs-and-Flexible-Normalization.pdf 06316-GAN-Models-from-an-Energy-Perspective-I-GAN-Digging-Holes-Jumping-into-Holes.pdf 06331-GAN-Models-from-an-Energy-Perspective-II-GAN-Analysis-Sampling.pdf 06377-Appreciation-of-the-Identity-A-Tr-A.pdf 06387-Clever-Gradient-Cutting-Implementing-GAN-Models-with-a-Single-Loss.pdf 06394-A-Brief-Introduction-to-the-Non-Adversarial-Generative-Model-GLANN.pdf 06407-Constructing-an-Explicit-Always-Invertible-Matrix.pdf 06409-O-GAN-A-Simple-Modification-to-Turn-the-GAN-Discriminator-into-an-Encoder.pdf 06418-Make-Keras-Cooler-Layer-wise-Learning-Rates-and-Free-Gradients.pdf 06482-Continuous-Flow-Invertible-ResNet-The-Ultimate-Brute-Force-Aesthetics.pdf 06534-Sharing-Drawing-a-3D-Order-3-Magic-Cube-with-LaTeX-MathJax.pdf 06540-Sharing-an-Unsupervised-Mining-of-Domain-Specific-Vocabulary.pdf 06549-From-DCGAN-to-SELF-MOD-An-Overview-of-GAN-Architecture-Evolution.pdf 06575-Making-Keras-Cooler-Intermediate-Variables-Weight-Averaging-and-Safe-Generators.pdf 06583-Optimization-Algorithms-from-a-Dynamics-Perspective-IV-The-Third-Stage-of-GAN.pdf 06612-GAN-Models-from-an-Energy-Perspective-III-Generative-Model-Energy-Model.pdf 06620-A-Discussion-on-Function-Smoothing-Differentiable-Approximations-of-Non-differentiable-Functions.pdf 06621-ON-LSTM-Expressing-Hierarchical-Structures-with-Ordered-Neurons.pdf 06671-A-Lightweight-Information-Extraction-Model-Based-on-DGCNN-and-Probabilistic-Graphs.pdf 06705-A-Discussion-on-Reparameterization-From-Normal-Distribution-to-Gumbel-Softmax.pdf 06736-When-BERT-Meets-Keras-This-Might-Be-the-Easiest-Way-to-Use-BERT.pdf 06747-A-Brief-Discussion-on-Unbiased-and-Biased-Estimation.pdf 06760-A-Brief-Introduction-to-VQ-VAE-Vector-Quantized-AutoEncoder.pdf 06771-Bert-based-NL2SQL-Model-A-Concise-Baseline.pdf 06784-When-You-Jump-Rope-Have-You-Ever-Thought-About-the-Shape-of-the-Curve.pdf 06794-Trading-Time-for-Performance-Keras-Gradient-Accumulation-Optimizer.pdf 06810-Making-Keras-Cooler-Layer-in-Layer-and-Masking.pdf 06818-Reflection-Can-Two-Elliptical-Sheets-Be-Glued-into-a-3D-Solid.pdf 06853-Born-for-Efficiency-From-Standard-Attention-to-Sparse-Attention.pdf 06869-Keras-Implementation-of-Two-Optimizers-Lookahead-and-LazyOptimizer.pdf 06877-Bidirectional-Decoding-in-seq2seq.pdf 06906-Open-Sourcing-a-Version-of-the-DGCNN-Reading-Comprehension-Question-Answering-Model-Keras-Version.pdf 06910-Introduction-to-HSIC-An-Interesting-Idea-for-Determining-Correlation.pdf 06915-My-Own-Implementation-of-bert4keras.pdf 06919-Post-Competition-Notes-on-Baidu-Entity-Linking-Behavior-Modeling-and-Entity-Linking.pdf 06920-Rewriting-the-Previous-New-Word-Discovery-Algorithm-Faster-and-Better-New-Word-Discovery.pdf 06933-From-Language-Models-to-Seq2Seq-Transformer-is-All-About-the-Mask.pdf 06985-Make-Keras-Cooler-Layer-and-Model-Reuse-Techniques.pdf 06992-What-Role-Does-Batch-Normalization-Actually-Play-A-Behind-Closed-Doors-Analysis.pdf 07006-The-Minimum-Entropy-Principle-V-Step-by-Step-Community-Detection-and-Clustering.pdf 07031-When-Can-the-Speedup-Ratio-of-Multi-processing-Be-Greater-Than-1.pdf 07038-From-Denoising-Autoencoders-to-Generative-Models.pdf 07055-Keras-The-Gold-Standard-of-TensorFlow.pdf 07063-JoSE-Word-and-Sentence-Vectors-on-the-Sphere.pdf 07076-Distribution-of-the-Angle-Between-Two-Random-Vectors-in-n-Dimensional-Space.pdf 07094-A-Brief-Introduction-and-Implementation-of-6-Derived-Optimizers.pdf 07105-Cascading-Rejection-A-Simple-and-Effective-Way-to-Improve-GAN-Performance.pdf 07115-The-Universal-Seq2Seq-Reading-Comprehension-and-Question-Answering-Based-on-Seq2Seq.pdf 07124-Conditional-Text-Generation-Based-on-Conditional-Layer-Normalization.pdf 07148-Non-Autoregressive-is-Not-Bad-Reading-Comprehension-QA-Based-on-MLM.pdf 07161-Triple-Extraction-with-bert4keras.pdf 07169-Self-Orthogonality-Module-A-Plug-and-Play-Kernel-Orthogonalization-Module.pdf 07180-Understanding-Model-Parameter-Initialization-Strategies-from-a-Geometric-Perspective.pdf 07187-Breaking-Constraints-Enhancing-Models-Improving-ALBERT-Performance-with-One-Line-of-Code.pdf 07196-Your-CRF-Layer-s-Learning-Rate-Might-Not-Be-Large-Enough.pdf 07210-Designing-GANs-Another-GAN-Production-Workshop.pdf 07213-Used-CRF-Why-Not-Learn-About-the-Faster-MEMM.pdf 07234-A-Brief-Discussion-on-Adversarial-Training-Significance-Methods-and-Reflections-with-Keras-Implementation.pdf 07259-A-Brief-Analysis-and-Countermeasures-for-Exposure-Bias-in-Seq2Seq.pdf 07292-Now-You-Can-Play-with-Chinese-GPT-2-Using-Keras-GPT2-ML.pdf 07302-Analysis-of-the-AdaFactor-Optimizer-with-Open-Source-Implementation.pdf 07309-How-the-Two-Elementary-Function-Approximations-of-GELU-Were-Derived.pdf 07321-With-bert4keras-in-Hand-I-Have-the-Baselines-Baidu-LIC2020.pdf 07325-Breaking-the-Bottleneck-Building-a-More-Powerful-Transformer.pdf 07343-EAE-Autoencoder-BN-Maximum-Entropy-Generative-Model.pdf 07359-Generalizing-Softmax-Cross-Entropy-to-Multi-Label-Classification.pdf 07367-Recomputation-Technique-for-Saving-VRAM-Now-Has-a-Keras-Version.pdf 07381-Variational-Autoencoders-V-VAE-BN-A-Better-VAE.pdf 07387-A-Brief-Analysis-of-the-AdaX-Optimizer-with-Open-Source-Implementation.pdf 07388-From-EMD-and-WMD-to-WRD-Similarity-Calculation-of-Text-Vector-Sequences.pdf 07427-Having-Your-Cake-and-Eating-It-Too-SimBERT-Model-Integrating-Retrieval-and-Generation.pdf 07430-Google-s-New-Synthesizer-We-Don-t-Understand-Self-Attention-Well-Enough.pdf 07466-Random-Thoughts-on-Generalization-From-Random-Noise-and-Gradient-Penalty-to-Virtual-Adversarial-Training.pdf 07469-Why-Gradient-Clipping-Accelerates-Training-A-Brief-Analysis.pdf 07476-Unsupervised-Word-Segmentation-and-Syntactic-Analysis-So-BERT-Can-Be-Used-This-Way.pdf 07500-How-to-Deal-with-the-Never-Ending-Problem-in-Seq2Seq.pdf 07521-Optimization-from-a-Sampling-Perspective-A-Unified-View-of-Differentiable-and-Non-Differentiable-Optimization.pdf 07533-Integrated-Gradients-A-Novel-Neural-Network-Visualization-Method.pdf 07546-Exploration-of-Linear-Attention-Does-Attention-Have-to-Have-a-Softmax.pdf 07574-Powerful-NVAE-No-More-Saying-VAE-Generated-Images-are-Blurry.pdf 07575-BERT-of-Theseus-A-Model-Compression-Method-Based-on-Module-Replacement.pdf 07611-A-Few-Words-on-the-National-Youth-Science-and-Technology-Innovation-Contest.pdf 07615-Mitigating-Class-Imbalance-via-Mutual-Information.pdf 07630-BERT-that-Learns-to-Ask-End-to-End-Question-Answer-Pair-Construction-from-Passages.pdf 07643-Do-We-Really-Need-to-Reduce-the-Training-Loss-to-Zero.pdf 07661-Modifying-Transformer-Structure-to-Design-a-Faster-and-Better-MLM-Model.pdf 07681-Is-L2-Regularization-Not-as-Good-as-Imagined-It-Might-Be-Due-to-Weight-Scale-Shifting.pdf 07695-The-Principle-of-Minimum-Entropy-VI-How-to-Choose-the-Dimension-of-Word-Embeddings.pdf 07708-Revisiting-Class-Imbalance-Connections-Between-Weight-Adjustment-and-Modified-Loss-Functions.pdf 07718-Building-Your-Own-DialoGPT-A-Generative-Multi-turn-Dialogue-Model-Based-on-LM.pdf 07725-Variational-Autoencoders-Part-6-An-Attempt-to-Understand-VAE-from-a-Geometric-Perspective.pdf 07737-Policy-Gradient-and-Zeroth-Order-Optimization-Different-Paths-to-the-Same-Destination.pdf 07758-Speeding-Up-Without-Performance-Loss-Chinese-WoBERT-Based-on-Word-Granularity.pdf 07764-Is-GPT-3-Necessary-No-BERT-s-MLM-Model-Can-Also-Do-Few-Shot-Learning.pdf 07782-The-1000th-Article.pdf 07787-Optimization-Algorithms-from-a-Dynamics-Perspective-V-Why-Should-the-Learning-Rate-Not-Be-Too-Small.pdf 07805-How-to-Split-a-Validation-Set-Closer-to-the-Test-Set.pdf 07809-What-Grade-Level-Can-BERT-Reach-Seq2Seq-Directly-Tackling-Primary-School-Math-Word-Problems.pdf 07818-TeaForN-Making-Teacher-Forcing-a-Bit-More-Farsighted.pdf 07846-Before-Using-ALBERT-and-ELECTRA-Make-Sure-You-Really-Understand-Them.pdf 07867-That-Leaderboard-Topping-T5-Model-Can-Now-Be-Used-for-Chinese.pdf 07877-When-GPT-Meets-Chinese-Chess-From-Writing-Articles-and-Solving-Problems-to-Playing-a-Game.pdf 07888-Let-s-Talk-About-the-Gradient-Vanishing-Exploding-Problem-in-RNNs.pdf 07912-Playing-with-the-Currently-Largest-Chinese-GPT-2-Model-bert4keras.pdf 07919-The-Even-Order-Taylor-Expansion-of-x-at-x-0-is-Always-Positive.pdf 07921-Performer-Linearizing-Attention-Complexity-with-Random-Projections.pdf 07947-Hierarchical-Decomposition-of-Position-Embeddings-Enabling-BERT-to-Handle-Ultra-Long-Text.pdf 07961-Turtle-Fish-Journal-An-All-Ceramsite-Undergravel-Filter-UGF-Ecological-Tank.pdf 07980-Optimization-Algorithms-from-a-Dynamics-Perspective-VI-Why-Doesn-t-SimSiam-Degenerate.pdf 07991-Mitchell-Approximation-Turning-Multiplication-into-Addition-with-Error-No-More-Than-1-9.pdf 08009-Optimization-Algorithms-from-a-Dynamics-Perspective-VII-SGD-SVM.pdf 08027-RealFormer-Transferring-Residuals-to-the-Attention-Matrix.pdf 08046-SPACES-An-Extract-Generate-Framework-for-Long-Text-Summarization-A-Summary-of-the-CAIL-Challenge.pdf 08062-Text-by-Searching-1-From-Text-Generation-to-Search-Sampling.pdf 08069-You-May-Not-Need-BERT-flow-A-Linear-Transformation-Comparable-to-BERT-flow.pdf 08084-Text-by-Search-II-From-MCMC-to-Simulated-Annealing.pdf 08119-Text-from-Search-III-Text-Sampling-based-on-BERT.pdf 08128-A-Theoretical-Analysis-Attempt-of-the-Repetition-Phenomenon-in-Seq2Seq.pdf 08130-Transformer-Positional-Encodings-that-Rack-Researchers-Brains.pdf 08159-How-Does-a-Binarized-Word-Vector-Model-Relate-to-Fruit-Flies.pdf 08180-Nystr-o-mformer-A-Linearized-Attention-Scheme-Based-on-Matrix-Decomposition.pdf 08194-Text-from-Search-4-Sentence-Construction-via-Addition-Deletion-and-Modification.pdf 08209-T5-PEGASUS-Open-Sourcing-a-Chinese-Generative-Pre-trained-Model.pdf 08213-Short-Text-Matching-Baseline-An-Attempt-to-Use-Pre-trained-Models-on-Desensitized-Data.pdf 08231-The-Road-to-Transformer-Upgrade-1.-Tracing-the-Origins-of-Sinusoidal-Position-Encoding.pdf 08244-The-Success-of-WGAN-May-Have-Little-to-Do-with-Wasserstein-Distance.pdf 08265-The-Road-to-Transformer-Upgrade-2.-Rotary-Position-Embedding-RoPE.pdf 08295-P-tuning-Automatically-Constructing-Templates-to-Unleash-the-Potential-of-Language-Models.pdf 08321-Which-Unsupervised-Semantic-Similarity-Method-is-the-Best-A-Comprehensive-Evaluation.pdf 08337-Sohu-Text-Matching-A-Multi-task-Baseline-Based-on-Conditional-LayerNorm.pdf 08338-The-Road-to-Transformer-Upgrade-3.-From-Performer-to-Linear-Attention.pdf 08348-Is-it-still-SOTA-for-Chinese-tasks-We-added-some-experiments-to-SimCSE.pdf 08373-GlobalPointer-A-Unified-Way-to-Handle-Nested-and-Non-nested-NER.pdf 08397-The-Road-to-Transformer-Upgrade-4.-Rotary-Position-Embedding-for-2D-Positions.pdf 08404-Variational-Autoencoders-VII-VAE-on-the-Sphere-vMF-VAE.pdf 08431-A-Review-of-Some-Recent-Non-Transformer-Works.pdf 08444-Can-We-Losslessly-Enlarge-a-Transformer-Model-Part-1.pdf 08453-Orthogonal-Matrix-for-Transforming-One-Unit-Vector-to-Another.pdf 08454-SimBERTv2-is-Here-The-RoFormer-Sim-Model-Integrating-Retrieval-and-Generation.pdf 08471-Can-Contrastive-Learning-Use-Gradient-Accumulation.pdf 08475-UniVAE-A-Single-Model-Multi-Scale-VAE-Based-on-Transformer.pdf 08496-Dropout-Twice-Again-This-Time-It-Achieves-SOTA-in-Supervised-Tasks.pdf 08512-KL-Divergence-Bhattacharyya-Distance-and-Wasserstein-Distance-between-Two-Multivariate-Normal-Distributions.pdf 08541-Enhancing-RoFormer-Sim-with-Open-Source-Human-Annotated-Data.pdf 08578-Linear-Models-from-a-Probabilistic-Perspective-Does-Logistic-Regression-Have-an-Analytical-Solution.pdf 08586-FlatNCE-Is-the-Reason-for-Poor-Performance-of-Small-Batch-Contrastive-Learning-Actually-Floating-Point-Error.pdf 08601-The-Road-to-Transformer-Upgrade-5.-Linear-Attention-as-Infinite-Dimensions.pdf 08610-Linear-Transformers-are-Probably-Not-the-Models-You-Are-Waiting-For.pdf 08620-A-Brief-Discussion-on-Initialization-Parameterization-and-Normalization-of-Transformers.pdf 08634-Gradient-Accumulation-Hidden-in-Momentum-Fewer-Updates-Better-Results.pdf 08656-From-the-Triangle-Inequality-to-Margin-Softmax.pdf 08662-Global-Shuffling-of-Hundreds-of-GBs-of-Files-under-Limited-Memory-Python.pdf 08671-The-Once-Disliked-Pre-training-Task-NSP-Achieves-Excellent-Zero-Shot-Results.pdf 08679-The-Amazing-Johnson-Lindenstrauss-Lemma-Theory-Edition.pdf 08706-The-Amazing-Johnson-Lindenstrauss-Lemma-Applications.pdf 08711-Analysis-of-the-Usability-of-the-Dimension-Formula-n-8.33-N.pdf 08715-Doubts-and-Communication-Regarding-the-Originality-of-WhiteningBERT.pdf 08718-Constructing-Smooth-Approximations-of-Non-Smooth-Functions-Using-the-Dirac-Delta-Function.pdf 08725-Reflections-on-Dimension-Averaging-Strategies-for-Non-Square-Matrices-in-Initialization-Methods.pdf 08728-CAN-A-Simple-Post-processing-Trick-to-Improve-Classification-Performance-via-Prior-Distribution.pdf 08739-bert4keras-in-Hand-Baselines-I-Have-CLUE-Benchmark-Code.pdf 08747-A-Discussion-on-Model-Optimization-Why-is-the-Initial-Standard-Deviation-of-BERT-0.02.pdf 08757-A-New-WGAN-Scheme-Implementing-Lipschitz-Constraint-via-Gradient-Normalization.pdf 08764-ChildTuning-Try-Adding-Dropout-to-the-Gradients.pdf 08770-MLM-and-MAE-from-the-Perspective-of-Dropout-Some-New-Insights.pdf 08783-Nonsense-at-the-Start-Fabricated-Data-Truly-Angered-by-a-Miracle-Paper.pdf 08791-Variational-Autoencoders-Part-8-Estimating-Sample-Probability-Density.pdf 08796-An-Inequality-Between-Input-Gradient-Penalty-and-Parameter-Gradient-Penalty.pdf 08802-Seq2Seq-Prefix-Tree-A-New-Paradigm-for-Retrieval-Tasks-Taking-KgCLUE-as-an-Example.pdf 08823-Understanding-the-Scale-Operation-in-Attention-from-the-Perspective-of-Entropy-Invariance.pdf 08829-Entropy-Normalization-of-Probability-Distributions.pdf 08833-SquarePlus-Possibly-the-Simplest-Smooth-Approximation-of-ReLU.pdf 08847-CoSENT-I-A-More-Effective-Sentence-Embedding-Scheme-than-Sentence-BERT.pdf 08860-CoSENT-Part-2-How-Big-is-the-Gap-Between-Representation-based-and-Interaction-based-Matching.pdf 08870-Notes-on-Multi-Task-Learning-I-In-the-Name-of-Loss.pdf 08877-Efficient-GlobalPointer-Fewer-Parameters-Better-Performance.pdf 08888-GPLinker-Joint-Entity-and-Relation-Extraction-based-on-GlobalPointer.pdf 08896-Multi-Task-Learning-Notes-II-Acting-on-Gradients.pdf 08907-Multi-Task-Learning-Chat-III-Prioritizing-Primary-and-Secondary-Tasks.pdf 08926-GPLinker-Joint-Event-Extraction-Based-on-GlobalPointer.pdf 08934-FLASH-Possibly-the-Most-Interesting-Efficient-Transformer-Design-Recently.pdf 08968-Exponentiated-Gradient-Descent-Meta-Learning-Adaptive-Learning-Rate.pdf 08978-What-are-the-Difficulties-in-Training-a-1000-layer-Transformer.pdf 08990-Does-Gated-Attention-Unit-GAU-Still-Need-Warmup.pdf 08994-Why-Do-We-Need-Residuals-A-Perspective-from-DeepNet.pdf 08998-RoFormerV2-Exploring-the-Limits-of-Natural-Language-Understanding.pdf 09009-Why-is-Pre-Norm-Less-Effective-than-Post-Norm.pdf 09019-I-Heard-Attention-and-Softmax-Go-Better-Together.pdf 09034-A-Quick-Derivation-of-Entropy-Invariant-Softmax.pdf 09039-What-Should-KL-Divergence-Look-Like-Under-GlobalPointer.pdf 09046-Does-Your-Language-Model-Have-Unpredictable-Words.pdf 09052-GAU-A-First-Look-at-the-Fast-Effective-and-Efficient-Next-Generation-Attention.pdf 09059-Accelerating-Training-with-Mixed-Precision-and-XLA-in-bert4keras.pdf 09064-Soft-label-Version-of-Multi-label-Softmax-Cross-Entropy.pdf 09070-Several-Inequalities-for-the-LogSumExp-Operation.pdf 09079-When-BERT-whitening-Introduces-Hyperparameters-There-s-Always-One-for-You.pdf 09085-Constructing-Discrete-Probability-Distributions-from-the-Perspective-of-Reparameterization.pdf 09098-How-to-Train-Your-Accuracy.pdf 09105-A-Theoretical-Defect-and-Countermeasure-of-Transformers-with-Relative-Position-Encoding.pdf 09119-Talk-on-Generative-Diffusion-Models-I-DDPM-Demolition-Construction.pdf 09138-Ladder-Side-Tuning-A-Scaling-Ladder-for-Pre-trained-Models.pdf 09147-A-Brief-Analysis-of-the-Hubness-Phenomenon-in-the-Curse-of-Dimensionality.pdf 09152-Generative-Diffusion-Models-2-DDPM-Autoregressive-VAE.pdf 09158-An-Unsuccessful-Attempt-Generalizing-Multi-label-Cross-entropy-to-n-sets-of-m-class-classification.pdf 09164-Generative-Diffusion-Models-Part-3-DDPM-Bayes-Denoising.pdf 09181-Discussion-on-Generative-Diffusion-Models-Part-4-DDIM-High-Perspective-DDPM.pdf 09209-Generative-Diffusion-Models-Part-5-SDE-General-Framework.pdf 09228-Generative-Diffusion-Models-6-ODE-Perspective-of-the-General-Framework.pdf 09245-Notes-on-Generative-Diffusion-Models-7-Optimal-Diffusion-Variance-Estimation-Part-1.pdf 09246-Generative-Diffusion-Models-Part-8-Optimal-Diffusion-Variance-Estimation-II.pdf 09257-Talk-on-Generative-Diffusion-Models-9-Conditional-Control-of-Generation-Results.pdf 09262-Generative-Diffusion-Models-Part-10-Unified-Diffusion-Model-Theory.pdf 09271-Generative-Diffusion-Models-11-Unified-Diffusion-Model-Application.pdf 09280-Generative-Diffusion-Models-Part-12-Directly-Tackling-the-Diffusion-ODE.pdf 09291-A-Brief-Attempt-at-the-Cross-Combinatorial-Counting-Problem.pdf 09305-Generative-Diffusion-Models-13-From-Universal-Gravitation-to-Diffusion-Models.pdf 09324-Probability-of-n-Random-Points-in-a-Circle-Falling-Within-a-Sector-with-Central-Angle.pdf 09336-Accelerating-Retrieval-of-Interactive-Similarity-Models-Using-CUR-Decomposition.pdf 09341-CoSENT-Part-3-As-a-Loss-Function-for-Interaction-based-Similarity.pdf 09344-Some-Alchemy-Strategies-Derived-from-the-Amos-Optimizer.pdf 09359-Guiding-Self-Supervised-Learning-with-the-Heat-Equation.pdf 09365-A-Simple-Solution-for-Controlling-Xgimi-Projectors-with-Xiao-Ai-in-Smart-Homes.pdf 09368-From-Local-to-Global-Geodesic-Distance-for-Semantic-Similarity.pdf 09370-Generative-Diffusion-Models-14-General-Steps-for-Constructing-ODEs-Part-1.pdf 09379-Generative-Diffusion-Models-Part-15-General-Steps-for-Constructing-ODEs-Part-2.pdf 09403-Transformer-Upgrade-Road-6.-Completeness-Analysis-of-Rotary-Positional-Embedding.pdf 09405-Analysis-of-the-Principles-of-Zero-Cold-Water-Technology-for-Smart-Home-Water-Heaters.pdf 09431-Transformer-Upgrade-Road-7.-Length-Extrapolation-and-Local-Attention.pdf 09444-The-Road-to-Transformer-Upgrade-8.-Length-Extrapolation-and-Positional-Robustness.pdf 09461-Deriving-the-Continuity-Equation-and-Fokker-Planck-Equation-via-the-Test-Function-Method.pdf 09467-Generative-Diffusion-Models-Part-16-Wasserstein-Distance-Score-Matching.pdf 09473-Google-s-New-Discovered-Optimizer-Lion-A-Training-Lion-Combining-Efficiency-and-Effectiveness.pdf 09497-Generative-Diffusion-Models-17-General-Steps-for-Constructing-ODEs-Part-3.pdf 09509-Generative-Diffusion-Models-18-Score-Matching-Conditional-Score-Matching.pdf 09512-Tiger-An-Extremely-Stingy-Optimizer.pdf 09526-A-Simple-Solution-to-Mitigate-Overconfidence-in-Cross-Entropy.pdf 09529-Why-are-Current-LLMs-All-Decoder-only-Architectures.pdf 09547-FAQ-Why-are-current-LLMs-all-Decoder-only-architectures.pdf 09554-Google-s-New-Work-Attempts-to-Resurrect-RNN-Can-RNN-Shine-Again.pdf 09577-The-Magical-Role-of-Bias-RoPE-Bias-Better-Length-Extrapolation.pdf 09588-Entropy-Invariant-Attention-from-the-Perspective-of-the-JL-Lemma.pdf 09590-LoRA-from-a-Gradient-Perspective-Introduction-Analysis-Conjectures-and-Generalization.pdf 09593-Two-Interesting-Discoveries-about-Attention-and-Softmax-Robustness-and-Information-Content.pdf 09595-How-to-Measure-the-Sparsity-of-Data.pdf 09603-Transformer-Upgrade-Road-9.-A-New-Idea-for-Global-Length-Extrapolation.pdf 09607-Deriving-Model-Scaling-Laws-Based-on-the-Quantization-Hypothesis.pdf 09617-NBCE-Extending-the-Context-Length-of-LLMs-using-Naive-Bayes.pdf 09632-Some-Supplementary-Notes-and-Analysis-on-the-NBCE-Method.pdf 09648-Naive-Bayes-is-all-you-need.pdf 09660-Gradient-Flow-Exploring-the-Path-to-the-Minimum.pdf 09662-Generative-Diffusion-Models-19-GAN-as-a-Diffusion-ODE.pdf 09668-Generative-Diffusion-Models-Part-20-From-ReFlow-to-WGAN-GP.pdf 09675-Transformer-Upgrade-Road-10.-RoPE-is-a-ary-Encoding.pdf 09687-When-Generative-Models-Run-Amok-Will-the-Internet-Suffer-from-Mad-Cow-Disease.pdf 09698-Re-exploring-Shared-Embeddings-at-the-Output-of-Language-Models.pdf 09706-Transformer-Upgrade-Road-11.-Pushing-base-Positional-Encoding-to-the-Limit.pdf 09708-Transformer-Upgrade-Road-12.-ReRoPE-with-Infinite-Extrapolation.pdf 09728-Transformer-Upgrade-Road-13.-Inverse-Leaky-ReRoPE.pdf 09731-Transformer-Upgrade-Road-14.-When-HWFA-Meets-ReRoPE.pdf 09736-Embedding-Anomalies-and-Countermeasures-under-Lion-Tiger-Optimizer-Training.pdf 09752-BytePiece-A-Purer-Tokenizer-with-Higher-Compression-Ratio.pdf 09762-A-Problem-and-Countermeasure-for-Large-Vocabulary-Language-Models-in-Text-Continuation-Tasks.pdf 09768-A-Brief-Exploration-of-Stochastic-Tokenization-From-Viterbi-Decoding-to-Viterbi-Sampling.pdf 09775-Minimum-Value-of-a-b-c-for-N-ab-c-in-the-Set-of-Natural-Numbers.pdf 09783-Mind-Blowing-Can-Non-linear-RNNs-Actually-Be-Computed-in-Parallel.pdf 09787-Pre-train-it-and-Transformer-s-Long-Sequence-Performance-Can-Still-Improve-Significantly.pdf 09797-EMO-A-Classification-Loss-Function-Designed-Based-on-Optimal-Transport.pdf 09811-Random-Word-Segmentation-Revisited-From-Viterbi-Sampling-to-Perfect-Sampling-Algorithms.pdf 09812-Looking-at-the-Scale-Operation-of-Attention-from-the-Perspective-of-Gradient-Maximization.pdf 09826-Embarrassingly-Simple-FSQ-Rounding-Surpasses-VQ-VAE.pdf 09844-VQ-the-Key-and-Transformer-Complexity-Becomes-Linear.pdf 09855-Life-Notes-The-Ultimate-Destination-of-Woks-is-the-Iron-Pan.pdf 09859-Transformer-Upgrade-Road-15.-Key-Normalization-for-Length-Extrapolation.pdf 09862-I-Found-Traces-of-Transformer-VQ-in-Performer.pdf 09881-Generative-Diffusion-Models-Part-21-Accelerating-ODE-Sampling-with-the-Mean-Value-Theorem.pdf 09889-Can-the-Attention-Mechanism-Really-Focus-Attention.pdf 09902-Making-Alchemy-More-Scientific-I-Convergence-of-Average-Loss-in-SGD.pdf 09907-I-Developed-an-Auxiliary-Website-for-Browsing-Papers-Cool-Papers.pdf 09920-Happy-New-Year-Recording-the-Development-Experience-of-Cool-Papers.pdf 09931-If-Local-Cosine-Similarity-is-Large-is-Global-Cosine-Similarity-Necessarily-Large.pdf 09938-Unconventional-Ways-to-Make-Python-Retry-Code-More-Elegant.pdf 09948-The-Road-to-Transformer-Upgrade-16.-A-Retrospective-on-Length-Extrapolation-Techniques.pdf 09969-Idempotent-Generative-Network-IGN-A-GAN-Attempting-to-Merge-Discrimination-and-Generation-into-One.pdf 09978-A-More-Convenient-Way-to-Open-Cool-Papers-Chrome-Redirect-Extension.pdf 09984-Building-a-Car-Behind-Closed-Doors-A-Brief-Discussion-on-Multimodal-Ideas-I-Lossless-Input.pdf 10001-Can-LoRA-Improve-Further-by-Configuring-Different-Learning-Rates.pdf 10007-Fitting-One-Dimensional-Probability-Density-Functions-with-Fourier-Series.pdf 10017-Chapter-of-Space-and-Time-Viewing-Attention-as-an-RNN-with-Quadratic-Complexity.pdf 10040-Transformer-Upgrade-Road-17.-Simple-Thoughts-on-Multimodal-Position-Encodings.pdf 10047-Generative-Diffusion-Models-22-Signal-to-Noise-Ratio-and-Large-Image-Generation-Part-1.pdf 10055-Casual-Talk-on-Generative-Diffusion-Models-23-Signal-to-Noise-Ratio-and-High-Resolution-Generation-Part-2.pdf 10077-Generative-Diffusion-Models-Part-24-Taking-Fewer-Shortcuts-to-Arrive-Faster.pdf 10085-Generative-Diffusion-Model-Chat-25-Identity-Based-Distillation-Part-1.pdf 10088-Cool-Papers-Update-Building-a-Simple-In-Site-Search-System.pdf 10091-The-Tug-of-War-Between-Cache-and-Performance-From-MHA-MQA-GQA-to-MLA.pdf 10114-Revisiting-SSM-I-Linear-Systems-and-HiPPO-Matrices.pdf 10122-The-Road-to-Transformer-Upgrade-18.-Principles-for-Selecting-the-RoPE-Base.pdf 10137-Revisiting-SSM-Part-2-Some-Legacy-Issues-of-HiPPO.pdf 10145-The-Road-to-Probability-Distributions-A-Review-of-Softmax-and-Its-Alternatives.pdf 10162-Revisiting-SSM-Part-3-Efficient-Computation-of-HiPPO-S4.pdf 10180-Revisiting-SSM-IV-A-New-Perspective-from-Rational-Generating-Functions.pdf 10197-Private-Reflections-on-Multimodal-Ideas-Part-2-Autoregression.pdf 10226-Aligning-with-Full-Fine-Tuning-This-is-the-Most-Brilliant-LoRA-Improvement-I-ve-Seen-Part-1.pdf 10240-Life-Notes-Cooking-Rice-Water-with-an-Electric-Rice-Cooker.pdf 10249-Monarch-Matrices-Computationally-Efficient-Sparse-Matrix-Decomposition.pdf 10266-Aligning-with-Full-Fine-Tuning-This-is-the-Most-Brilliant-LoRA-Improvement-I-ve-Seen-Part-2.pdf 10289-The-Road-to-Optimal-Distribution-Minimization-in-Probability-Space.pdf 10311-Some-New-Attempts-at-Cool-Papers-Site-Search.pdf 10320-Making-MathJax-Better-Compatible-with-Google-Translate-and-Lazy-Loading.pdf 10332-A-Near-Perfect-Solution-to-the-Conflict-Between-MathJax-and-Marked.pdf 10347-Why-Do-Decoder-only-LLMs-Need-Positional-Encoding.pdf 10352-Building-a-Car-Behind-Closed-Doors-A-Brief-Discussion-on-Multimodal-Ideas-Part-3-Position-Encoding.pdf 10366-The-Road-to-Low-Rank-Approximation-Part-1-Pseudo-inverse.pdf 10373-Softmax-Sequel-Finding-Smooth-Approximations-for-Top-K.pdf 10394-Implementing-Smart-Gas-Stove-Shut-off-Using-Flameout-Protection-Smart-Switch.pdf 10407-The-Road-to-Low-Rank-Approximation-Part-2-SVD.pdf 10427-The-Road-to-Low-Rank-Approximation-III-CR.pdf 10474-Making-MathJax-Formulas-Auto-Scale-with-Window-Size.pdf 10480-Cool-Papers-Browser-Extension-Upgraded-to-v0.2.0.pdf 10489-The-Rotation-Trick-for-VQ-A-General-Extension-of-Straight-Through-Estimation.pdf 10501-The-Road-to-Low-Rank-Approximation-IV-Interpolative-Decomposition-ID.pdf 10519-Another-VQ-Trick-Adding-a-Linear-Transformation-to-the-Codebook.pdf 10542-How-Should-the-Learning-Rate-Change-as-the-Batch-Size-Increases.pdf 10563-How-Does-Adam-s-Epsilon-Affect-the-Scaling-Law-of-Learning-Rate.pdf 10567-Generative-Diffusion-Models-Part-26-Identity-Based-Distillation-Part-2.pdf 10588-Adaptive-Learning-Rate-Optimizers-from-the-Perspective-of-Hessian-Approximation.pdf 10592-An-Appreciation-of-the-Muon-Optimizer-The-Essential-Leap-from-Vectors-to-Matrices.pdf 10617-Generative-Diffusion-Models-Part-27-Step-Size-as-Conditional-Input.pdf 10633-Generative-Diffusion-Models-Part-28-A-Step-by-Step-Understanding-of-Consistency-Models.pdf 10648-Reflections-on-Spectral-Norm-Gradients-and-a-New-Type-of-Weight-Decay.pdf 10657-Why-is-the-Default-Norm-for-Gradient-Clipping-1.pdf 10662-The-Road-to-Low-Rank-Approximation-V-CUR.pdf 10667-Flowing-with-Flow-TARFLOW-Flow-Models-Return-at-Full-Strength.pdf 10684-Intersection-Coordinates-of-Three-Spheres-Trilateration.pdf 10699-MoE-Tour-1.-From-a-Geometric-Perspective.pdf 10711-Generative-Diffusion-Models-Part-29-Discrete-Encoding-with-DDPM.pdf 10735-MoE-Tour-2.-Not-Worried-about-Scarcity-but-about-Inequality.pdf 10739-Muon-Sequel-Why-We-Chose-to-Try-Muon.pdf 10757-MoE-Grand-Tour-3.-A-Different-Approach-to-Allocation.pdf 10770-A-First-Look-at-MuP-Cross-Model-Scaling-Laws-for-Hyperparameter-Transfer.pdf 10795-Higher-Order-MuP-A-Simpler-Yet-Smarter-Spectral-Condition-Scaling.pdf 10815-MoE-Journey-4.-Invest-More-in-Difficult-Parts.pdf 10831-Finding-Alternatives-to-Normalization-via-Gradient-Approximation.pdf 10847-Effective-Rank-of-a-Matrix.pdf 10862-Transformer-Upgrade-Road-19-The-Second-Type-of-Rotary-Positional-Encoding.pdf 10869-Smart-Home-DIY-a-Zero-Cold-Water-System-Compatible-with-Mi-Home.pdf 10878-The-Derivative-of-SVD.pdf 10902-A-Probability-Inequality-Stare-at-it-until-it-becomes-obvious.pdf 10907-The-Road-to-Transformer-Upgrades-20.-Why-is-MLA-So-Good-Part-1.pdf 10922-Newton-Schulz-Iteration-for-the-msign-Operator-Part-1.pdf 10945-MoE-Tour-5.-Reflections-on-Uniform-Distribution.pdf 10958-Generative-Diffusion-Model-Chat-30-From-Instantaneous-Velocity-to-Average-Velocity.pdf 10972-Equioscillation-Theorem-Necessary-and-Sufficient-Conditions-for-Optimal-Polynomial-Approximation.pdf 10996-Newton-Schulz-Iteration-for-the-msign-Operator-Part-2.pdf 11006-Computing-Singular-Value-Clipping-mclip-via-msign-Part-1.pdf 11025-The-Derivative-of-the-msign-Operator.pdf 11033-A-Brief-History-of-Linear-Attention-From-Imitation-and-Innovation-to-Reciprocation.pdf 11056-What-Can-the-Matrix-Sign-Function-mcsgn-Calculate.pdf 11059-Calculating-Singular-Value-Clipping-mclip-via-msign-Part-2.pdf 11072-Efficient-Inversion-Method-for-Diagonal-Low-Rank-Triangular-Matrices.pdf 11111-The-Road-to-Transformer-Upgrade-21.-Why-is-MLA-so-Good-Part-2.pdf 11126-QK-Clip-Taking-Muon-a-Step-Further-on-the-Road-to-Scaling-Up.pdf 11158-Efficient-Computation-of-Matrix-Square-Roots-and-Inverse-Square-Roots.pdf 11175-Efficient-Computation-of-Matrix-r-th-Roots-and-Inverse-r-th-Roots.pdf 11196-Steepest-Descent-on-Manifolds-1.-SGD-Hypersphere.pdf 11206-Building-a-Portable-Side-Router-Based-on-Raspberry-Pi-Zero-2W.pdf 11215-Steepest-Descent-on-Manifolds-2.-Muon-Orthogonality.pdf 11221-Steepest-Descent-on-Manifolds-3.-Muon-Stiefel.pdf 11233-An-Identity-for-ReLU-GeLU-Swish.pdf 11241-Steepest-Descent-on-Manifolds-4.-Muon-Spectral-Sphere.pdf 11250-Cool-Papers-Update-Simple-Adaptation-for-Zotero-Connector.pdf 11260-Rethinking-Learning-Rate-and-Batch-Size-I-Current-Status.pdf 11267-Why-is-Adam-s-Update-RMS-0.2.pdf 11280-Rethinking-Learning-Rate-and-Batch-Size-Part-2-Mean-Field.pdf 11285-Rethinking-Learning-Rate-and-Batch-Size-Part-3-Muon.pdf 11301-Rethinking-Learning-Rate-and-Batch-Size-Part-IV-EMA.pdf 11307-Asymptotic-Estimation-of-Weight-RMS-for-AdamW.pdf 11320-Why-do-Linear-Attention-Models-Add-Short-Conv.pdf 11328-DiVeQ-A-Very-Concise-VQ-Training-Scheme.pdf 11335-Fast-Estimation-of-the-Spectral-Norm-of-Random-Matrices.pdf 11340-Beyond-MuP-1.-Three-Characteristics-of-a-Good-Model.pdf 11371-Low-Precision-Attention-May-Have-Biased-Rounding-Errors.pdf 11388-Steepest-Descent-on-Manifolds-5.-Dual-Gradient-Descent.pdf 11390-Asymptotic-Estimation-of-the-Maximum-of-n-Normal-Random-Variables.pdf 11404-Asymptotic-Estimation-of-Weight-RMS-for-AdamW-Part-2.pdf 11416-Muon-Optimizer-Guide-Quick-Start-and-Key-Details.pdf 11428-Generative-Diffusion-Models-31-Predicting-Data-Instead-of-Noise.pdf 11459-Weight-Decay-and-Learning-Rate-from-the-Perspective-of-Moving-Average.pdf 11469-Making-Alchemy-More-Scientific-II-Generalizing-Conclusions-to-Unbounded-Domains.pdf 11480-Making-Alchemy-More-Scientific-III-Last-Iterate-Loss-Convergence-of-SGD.pdf 11486-Why-does-DeltaNet-need-L2-Normalization.pdf 11494-Making-Alchemy-More-Scientific-IV-New-Identity-New-Learning-Rate.pdf list.txt