Trendshift - Ask AI

base on High-efficiency floating-point neural network inference operators for mobile, server, and Web # XNNPACK XNNPACK is a highly optimized solution for neural network inference on ARM, x86, WebAssembly, and RISC-V platforms. XNNPACK is not intended for direct use by deep learning practitioners and researchers; instead it provides low-level performance primitives for accelerating high-level machine learning frameworks, such as [TensorFlow Lite](https://www.tensorflow.org/lite), [TensorFlow.js](https://www.tensorflow.org/js), [PyTorch](https://pytorch.org/), [ONNX Runtime](https://onnxruntime.ai), [ExecuTorch](https://pytorch.org/executorch-overview), and [MediaPipe](https://mediapipe.dev). ## Supported Architectures - ARM64 on Android, iOS, macOS, Linux, and Windows - ARMv7 (with NEON) on Android - ARMv6 (with VFPv2) on Linux - x86 and x86-64 (up to AVX512) on Windows, Linux, macOS, Android, and iOS simulator - WebAssembly MVP - WebAssembly SIMD - [WebAssembly Relaxed SIMD](https://github.com/WebAssembly/relaxed-simd) (experimental) - RISC-V (RV32GC and RV64GC) ## Operator Coverage XNNPACK implements the following neural network operators: - 2D Convolution (including grouped and depthwise) - 2D Deconvolution (AKA Transposed Convolution) - 2D Average Pooling - 2D Max Pooling - 2D ArgMax Pooling (Max Pooling + indices) - 2D Unpooling - 2D Bilinear Resize - 2D Depth-to-Space (AKA Pixel Shuffle) - Add (including broadcasting, two inputs only) - Subtract (including broadcasting) - Divide (including broadcasting) - Maximum (including broadcasting) - Minimum (including broadcasting) - Multiply (including broadcasting) - Squared Difference (including broadcasting) - Global Average Pooling - Channel Shuffle - Fully Connected - Abs (absolute value) - Bankers' Rounding (rounding to nearest, ties to even) - Ceiling (rounding to integer above) - Clamp (includes ReLU and ReLU6) - Convert (includes fixed-point and half-precision quantization and dequantization) - Copy - ELU - Floor (rounding to integer below) - HardSwish - Leaky ReLU - Negate - Sigmoid - Softmax - Square - Tanh - Transpose - Truncation (rounding to integer towards zero) - PReLU All operators in XNNPACK support NHWC layout, but additionally allow custom stride along the **C**hannel dimension. Thus, operators can consume a subset of channels in the input tensor, and produce a subset of channels in the output tensor, providing a zero-cost Channel Split and Channel Concatenation operations. ## Performance ### Mobile phones The table below presents **single-threaded** performance of XNNPACK library on three generations of MobileNet models and three generations of Pixel phones. | Model | Pixel, ms | Pixel 2, ms | Pixel 3a, ms | | ----------------------- | :-------: | :---------: | :----------: | | FP32 MobileNet v1 1.0X | 82 | 86 | 88 | | FP32 MobileNet v2 1.0X | 49 | 53 | 55 | | FP32 MobileNet v3 Large | 39 | 42 | 44 | | FP32 MobileNet v3 Small | 12 | 14 | 14 | The following table presents **multi-threaded** (using as many threads as there are big cores) performance of XNNPACK library on three generations of MobileNet models and three generations of Pixel phones. | Model | Pixel, ms | Pixel 2, ms | Pixel 3a, ms | | ----------------------- | :-------: | :---------: | :----------: | | FP32 MobileNet v1 1.0X | 43 | 27 | 46 | | FP32 MobileNet v2 1.0X | 26 | 18 | 28 | | FP32 MobileNet v3 Large | 22 | 16 | 24 | | FP32 MobileNet v3 Small | 7 | 6 | 8 | Benchmarked on March 27, 2020 with `end2end_bench --benchmark_min_time=5` on an Android/ARM64 build with Android NDK r21 (`bazel build -c opt --config android_arm64 :end2end_bench`) and neural network models with randomized weights and inputs. ### Raspberry Pi The table below presents **multi-threaded** performance of XNNPACK library on three generations of MobileNet models and three generations of Raspberry Pi boards. | Model | RPi Zero W (BCM2835), ms | RPi 2 (BCM2836), ms | RPi 3+ (BCM2837B0), ms | RPi 4 (BCM2711), ms | RPi 4 (BCM2711, ARM64), ms | | ----------------------- | :----------------------: | :-----------------: | :--------------------: | :-----------------: | :------------------------: | | FP32 MobileNet v1 1.0X | 3919 | 302 | 114 | 72 | 77 | | FP32 MobileNet v2 1.0X | 1987 | 191 | 79 | 41 | 46 | | FP32 MobileNet v3 Large | 1658 | 161 | 67 | 38 | 40 | | FP32 MobileNet v3 Small | 474 | 50 | 22 | 13 | 15 | | INT8 MobileNet v1 1.0X | 2589 | 128 | 46 | 29 | 24 | | INT8 MobileNet v2 1.0X | 1495 | 82 | 30 | 20 | 17 | Benchmarked on Feb 8, 2022 with `end2end-bench --benchmark_min_time=5` on a Raspbian Buster build with CMake (`./scripts/build-local.sh`) and neural network models with randomized weights and inputs. INT8 inference was evaluated on per-channel quantization schema. ## Minimum build requirements - C11 - C++14 - Python 3 ## Publications - Marat Dukhan "The Indirect Convolution Algorithm". Presented on [Efficient Deep Learning for Compute Vision (ECV) 2019](https://sites.google.com/corp/view/ecv2019/) workshop ([slides](https://drive.google.com/file/d/1ZayB3By5ZxxQIRtN7UDq_JvPg1IYd3Ac/view), [paper on ArXiv](https://arxiv.org/abs/1907.02129)). - Erich Elsen, Marat Dukhan, Trevor Gale, Karen Simonyan "Fast Sparse ConvNets". [Paper on ArXiv](https://arxiv.org/abs/1911.09723), [pre-trained sparse models](https://github.com/google-research/google-research/tree/master/fastconvnets). - Marat Dukhan, Artsiom Ablavatski "The Two-Pass Softmax Algorithm". [Paper on ArXiv](https://arxiv.org/abs/2001.04438). - Yury Pisarchyk, Juhyun Lee "Efficient Memory Management for Deep Neural Net Inference". [Paper on ArXiv](https://arxiv.org/abs/2001.03288). ## Ecosystem ### Machine Learning Frameworks - [TensorFlow Lite](https://blog.tensorflow.org/2020/07/accelerating-tensorflow-lite-xnnpack-integration.html). - [TensorFlow.js WebAssembly backend](https://blog.tensorflow.org/2020/03/introducing-webassembly-backend-for-tensorflow-js.html). - [PyTorch Mobile](https://pytorch.org/mobile). - [ONNX Runtime Mobile](https://onnxruntime.ai/docs/execution-providers/Xnnpack-ExecutionProvider.html) - [MediaPipe for the Web](https://developers.googleblog.com/2020/01/mediapipe-on-web.html). - [Alibaba HALO (Heterogeneity-Aware Lowering and Optimization)](https://github.com/alibaba/heterogeneity-aware-lowering-and-optimization) - [Samsung ONE (On-device Neural Engine)](https://github.com/Samsung/ONE) ## Acknowledgements XNNPACK is based on [QNNPACK](https://github.com/pytorch/QNNPACK) library. Over time its codebase diverged a lot, and XNNPACK API is no longer compatible with QNNPACK. ", Assign "at most 3 tags" to the expected json: {"id":"1620","tags":[]} "only from the tags list I provide: [{"id":77,"name":"3d"},{"id":89,"name":"agent"},{"id":17,"name":"ai"},{"id":54,"name":"algorithm"},{"id":24,"name":"api"},{"id":44,"name":"authentication"},{"id":3,"name":"aws"},{"id":27,"name":"backend"},{"id":60,"name":"benchmark"},{"id":72,"name":"best-practices"},{"id":39,"name":"bitcoin"},{"id":37,"name":"blockchain"},{"id":1,"name":"blog"},{"id":45,"name":"bundler"},{"id":58,"name":"cache"},{"id":21,"name":"chat"},{"id":49,"name":"cicd"},{"id":4,"name":"cli"},{"id":64,"name":"cloud-native"},{"id":48,"name":"cms"},{"id":61,"name":"compiler"},{"id":68,"name":"containerization"},{"id":92,"name":"crm"},{"id":34,"name":"data"},{"id":47,"name":"database"},{"id":8,"name":"declarative-gui "},{"id":9,"name":"deploy-tool"},{"id":53,"name":"desktop-app"},{"id":6,"name":"dev-exp-lib"},{"id":59,"name":"dev-tool"},{"id":13,"name":"ecommerce"},{"id":26,"name":"editor"},{"id":66,"name":"emulator"},{"id":62,"name":"filesystem"},{"id":80,"name":"finance"},{"id":15,"name":"firmware"},{"id":73,"name":"for-fun"},{"id":2,"name":"framework"},{"id":11,"name":"frontend"},{"id":22,"name":"game"},{"id":81,"name":"game-engine "},{"id":23,"name":"graphql"},{"id":84,"name":"gui"},{"id":91,"name":"http"},{"id":5,"name":"http-client"},{"id":51,"name":"iac"},{"id":30,"name":"ide"},{"id":78,"name":"iot"},{"id":40,"name":"json"},{"id":83,"name":"julian"},{"id":38,"name":"k8s"},{"id":31,"name":"language"},{"id":10,"name":"learning-resource"},{"id":33,"name":"lib"},{"id":41,"name":"linter"},{"id":28,"name":"lms"},{"id":16,"name":"logging"},{"id":76,"name":"low-code"},{"id":90,"name":"message-queue"},{"id":42,"name":"mobile-app"},{"id":18,"name":"monitoring"},{"id":36,"name":"networking"},{"id":7,"name":"node-version"},{"id":55,"name":"nosql"},{"id":57,"name":"observability"},{"id":46,"name":"orm"},{"id":52,"name":"os"},{"id":14,"name":"parser"},{"id":74,"name":"react"},{"id":82,"name":"real-time"},{"id":56,"name":"robot"},{"id":65,"name":"runtime"},{"id":32,"name":"sdk"},{"id":71,"name":"search"},{"id":63,"name":"secrets"},{"id":25,"name":"security"},{"id":85,"name":"server"},{"id":86,"name":"serverless"},{"id":70,"name":"storage"},{"id":75,"name":"system-design"},{"id":79,"name":"terminal"},{"id":29,"name":"testing"},{"id":12,"name":"ui"},{"id":50,"name":"ux"},{"id":88,"name":"video"},{"id":20,"name":"web-app"},{"id":35,"name":"web-server"},{"id":43,"name":"webassembly"},{"id":69,"name":"workflow"},{"id":87,"name":"yaml"}]" returns me the "expected json"

AI prompts