Commit c479922ac for llama.cpp
commit c479922ac520a08969b4c1dc154d7bbb3c386d85
Author: Max Krasnyansky <maxk@qti.qualcomm.com>
Date: Tue Oct 6 15:24:17 2026 -0700
hexagon: CPY/CONCAT/CONT/DUP overhaul to use DMA/HVX for all cases (#30067)
* hex-cpy: replace more paths with dma and simplify l2flush
* hex-cpy: use DMA in all sametype paths
* hex-cpy: rewrite the rest of the copy paths (diff type) to use dma
* hex-concat: use dma for multi-dev path which also removes the need for l2-line alignment
* hex-concat: proper support for mdev splitting
* hex-concat: cleanup ctx and kern params usage
* hex-cpy: cleanup contex and remove left-over non-dma checks
* hex-cpy: clean dma_cpy naming
* hex-cpy: proper kernel params and kernel selection
* hex-build: resolve left-over rebase conflicts
* hex-cpy: update dev guide to clarify 128 byte alignment requirement
* hex-concat: make sure we go through mdev barrier
* hex-concat: make sure to flush dma-queue
* hex-cpy/concat: cleanup kparams and vtcm layout handling
* hex-cpy: remove dead check for contig (routed to diff kernel) and update comments
* hex-dev: update developer guide based on latest changes
* hex-cpy: safe skip of noop copies
* hex-dup: route DUP to CPY
diff --git a/docs/backend/snapdragon/developer.md b/docs/backend/snapdragon/developer.md
index 378d47653..3608250dc 100644
--- a/docs/backend/snapdragon/developer.md
+++ b/docs/backend/snapdragon/developer.md
@@ -37,9 +37,10 @@ In llama.cpp/GGML, each Hexagon session is mapped to a single GGML backend devic
`GGML_HEXAGON_DEVICES`, or `HTP0`, `HTP1` in legacy mode).
To support running models larger than 3.5GB on a single device, the Hexagon backend dynamically maps and unmaps buffers:
-- Buffers are allocated in shared DDR (RPCMEM) via file descriptors (`fastrpc_mmap` using `FASTRPC_MAP_FD_DELAYED`).
+- Buffers are allocated in shared DDR (RPCMEM) and mapped through FastRPC file descriptors. Non-pinned buffers use delayed
+ mappings (`FASTRPC_MAP_FD_DELAYED` or `FASTRPC_MAP_FD_DELAYED_EXTENDED`).
- Pinned buffers (such as KV cache and active compute buffers) remain mapped throughout execution.
-- Inactive weight buffers are dynamically mapped into the NPU session via `HAP_mmap()` during batch buffer preparation
+- Inactive weight buffers are dynamically mapped into the NPU session during batch buffer preparation
(`prep_op_bufs()` in `htp/main.c`) and unmapped via `htp_iface_munmap()` when no longer needed by the active batch.
- This dynamic sliding window allows a single NPU session to execute models that exceed the 3.5GB window.
@@ -55,6 +56,9 @@ Writing high-performance operators for Hexagon requires following specific guide
- Strongly prefer the `DDR -> DMA -> VTCM -> compute (HVX/HMX) -> VTCM -> DMA -> DDR` data flow.
- Direct HVX reads/writes from/to DDR are less efficient and should only be used as a fallback.
+- Use `dma_addr_t` only for DMA base and final addresses. Form a final address by adding a 32-bit byte offset to a
+ `dma_addr_t` tensor base address. This permits a 64-bit mapped base address on newer platforms while retaining 32-bit
+ relative addressing.
- The DMA queue is a strict FIFO where operations must be pushed and popped in strict order.
- Follow the pipelined multi-buffering sequence properly (typically 2x to 16x buffering) so every push has a corresponding pop:
@@ -66,7 +70,7 @@ Writing high-performance operators for Hexagon requires following specific guide
- Because every push must be matched by a pop, `dma_queue_flush()` is not required when the pipeline sequence is followed
properly. Flushing is only used in rare exceptions where a batch of operations is pushed without individual pops.
- Use the DMA queue interface from [`dma-queue.h`](../../../ggml/src/ggml-hexagon/htp/dma-queue.h)
- (`dma_queue_push_ddr_to_vtcm`, `dma_queue_pop`, `dma_queue_push_vtcm_to_ddr`).
+ (`dma_queue_push()`, `dma_queue_pop()`, and `dma_queue_flush()`).
See [`cumsum-ops.c`](../../../ggml/src/ggml-hexagon/htp/cumsum-ops.c) and
[`act-ops.c`](../../../ggml/src/ggml-hexagon/htp/act-ops.c) for reference implementations.
@@ -125,7 +129,6 @@ Writing high-performance operators for Hexagon requires following specific guide
- Do not add defensive NULL checks or assertions for internal framework pointers or required graph operands and outputs.
Internal pointers include `ctx`, `octx`, local context structs like `*ctx`, `kparams`, and worker callback `data`.
- These pointers are architectural invariants during kernel execution and host-side graph preparation.
- Graph compute receives allocated nodes with valid required `node->src[N]` and `node->data` pointers.
- Do not turn an invariant violation into an unsupported operation or missed fusion.
Checks such as `if (!octx || !octx->ctx)` clutter the code, obscure intent, and hide upstream errors.
- **Distinction**: `octx->src[N]` pointers *can* be NULL by design and must be checked when optional.
@@ -177,26 +180,28 @@ sessions.
- Shared tensor buffers reside in DDR (RPCMEM) with a 128-byte cache line granularity
(`HEX_L2_LINE_SIZE` = 128 bytes, `HTP_TENSOR_MDEV_LINE_SIZE`).
-- **Rule**: Multi-device work partitions must align destination write regions to 128-byte cache line boundaries so distinct
- devices never share or overwrite the same cache line.
+- **Rule**: Multi-device work partitions that write directly to DDR through HVX/L2 must align destination write regions to
+ 128-byte cache line boundaries so distinct devices never share or overwrite the same cache line.
+- DMA writes to DDR are not subject to this cache-line ownership rule. They may use smaller non-overlapping destination
+ ranges when the operator only writes through DMA.
### Partitioning Helpers in `htp-tensor.h`
Common partitioning logic is factored into reusable inline helpers in
[`htp-tensor.h`](../../../ggml/src/ggml-hexagon/htp/htp-tensor.h):
-1. [`htp_tensor_mdev_rows_per_chunk`](../../../ggml/src/ggml-hexagon/htp/htp-tensor.h#L67):
+1. [`htp_tensor_mdev_rows_per_chunk`](../../../ggml/src/ggml-hexagon/htp/htp-tensor.h#L71):
Determines the minimum number of rows per chunk so that the chunk byte size is a multiple of 128 bytes:
```
rows_per_chunk = 128 / hex_gcd_u32(row_size, 128)
```
- If row stride `nb[1]` is already a multiple of 128 bytes, `rows_per_chunk = 1`.
+ If the active row and outer strides are already multiples of 128 bytes, `rows_per_chunk = 1`.
Returns `false` if the tensor cannot be safely row-partitioned (such as unaligned base pointer, permuted layout,
or non-128-byte aligned outer strides).
-2. [`htp_tensor_mdev_partition`](../../../ggml/src/ggml-hexagon/htp/htp-tensor.h#L94):
+2. [`htp_tensor_mdev_partition`](../../../ggml/src/ggml-hexagon/htp/htp-tensor.h#L98):
Calculates the per-device work range `struct htp_tensor_mdev_range { uint32_t start; uint32_t count; }` given
`total_units`, `units_per_chunk`, `mdev_idx`, `mdev_count`, and the precomputed `mdev_count_div`.
Handles chunk distribution across devices, assigns remainder units to the last device, and automatically triggers
@@ -204,11 +209,10 @@ Common partitioning logic is factored into reusable inline helpers in
### Row-Partitioned Operators
-For row-wise operators
+For row-wise operators that write directly to DDR
(such as activations in [`act-ops.c`](../../../ggml/src/ggml-hexagon/htp/act-ops.c),
binary ops in [`binary-ops.c`](../../../ggml/src/ggml-hexagon/htp/binary-ops.c),
-unary ops in [`unary-ops.c`](../../../ggml/src/ggml-hexagon/htp/unary-ops.c), and
-sameshape copies in [`cpy-ops.c`](../../../ggml/src/ggml-hexagon/htp/cpy-ops.c)):
+and unary ops in [`unary-ops.c`](../../../ggml/src/ggml-hexagon/htp/unary-ops.c)):
```c
const uint32_t total_rows = ne01 * ne02 * ne03;
@@ -233,20 +237,19 @@ if (nrows == 0) {
### Element-Partitioned Operators
-For flat element-wise operations (such as reshape copies in
-[`cpy-ops.c`](../../../ggml/src/ggml-hexagon/htp/cpy-ops.c)):
+For flat element-wise operations that write directly to DDR:
- Partition total linear elements N = ne0 * ne1 * ne2 * ne3 in 128-byte cache line chunks (`elems_per_line = (elem_size == 4) ? 32 : 64`).
- Requires strict 1D contiguity:
- [`htp_tensor_is_contiguous(dst, elem_size)`](../../../ggml/src/ggml-hexagon/htp/htp-tensor.h#L28)
+ [`htp_tensor_is_contiguous(dst, elem_size)`](../../../ggml/src/ggml-hexagon/htp/htp-tensor.h#L32)
and 128-byte aligned destination pointer
- [`htp_tensor_mdev_data_aligned(dst)`](../../../ggml/src/ggml-hexagon/htp/htp-tensor.h#L47).
+ [`htp_tensor_mdev_data_aligned(dst)`](../../../ggml/src/ggml-hexagon/htp/htp-tensor.h#L51).
- If contiguous and aligned, pass `elems_per_line` to
- [`htp_tensor_mdev_partition`](../../../ggml/src/ggml-hexagon/htp/htp-tensor.h#L94);
+ [`htp_tensor_mdev_partition`](../../../ggml/src/ggml-hexagon/htp/htp-tensor.h#L98);
otherwise pass 0 to trigger Device 0 fallback.
### Single-Device Fallback (Device 0)
-- Fallback to Device 0 (`mdev.idx == 0`) when partitioning would cause cache line tearing or when work cannot be evenly distributed.
+- Fallback to Device 0 (`mdev.idx == 0`) when partitioning would cause cache line tearing or there are too few aligned chunks.
- Triggers:
1. Destination tensor cannot be safely partitioned (`rows_per_chunk == 0` or non-contiguous/unaligned buffer).
2. Total aligned chunks < `mdev_count`.
@@ -303,7 +306,7 @@ Multi-device execution synchronizes worker sessions through atomic fence slots a
(Input Prep) (Input Prep)
| |
Pre-Op Barrier ----------------------------- Pre-Op Barrier
- (mdev_sync_fence) (mdev_sync_fence)
+ (htp_mdev_group_barrier) (htp_mdev_group_barrier)
| |
Kernel Execution Kernel Execution
(Output Slice 0) (Output Slice 1)
@@ -326,10 +329,10 @@ Multi-device execution synchronizes worker sessions through atomic fence slots a
atomic_uint * my_fence = htp_mdev_fence_slot(fence_base, mdev_idx);
```
-- **Writing to fence ([`htp_fence_write`](../../../ggml/src/ggml-hexagon/htp/htp-fence.h#L18))**:
+- **Writing to fence ([`htp_fence_write`](../../../ggml/src/ggml-hexagon/htp/htp-fence.h#L17))**:
Stores `seq` and `status`, issues a `syncht` thread synchronization barrier, and flushes/invalidates the line
using `Q6_dccleaninva_A(fence)`.
-- **Reading from peer fence ([`htp_fence_read`](../../../ggml/src/ggml-hexagon/htp/htp-fence.h#L26))**:
+- **Reading from peer fence ([`htp_fence_read`](../../../ggml/src/ggml-hexagon/htp/htp-fence.h#L25))**:
Executes `Q6_dccleaninva_A(fence)` and `syncht` before reading atomic values to ensure fresh data from DDR.
### Deterministic Monotonic Sequence Numbers
@@ -348,7 +351,7 @@ Multi-device execution synchronizes worker sessions through atomic fence slots a
- In the kernel, ensure all pushed DMA operations have been popped in strict FIFO order to drain the queue.
- Use [`htp_tensor_flush_all()`](../../../ggml/src/ggml-hexagon/htp/htp-tensor.h) to flush specific dirty tensors back to DDR:
- - [`htp_tensor_flush_all()`](../../../ggml/src/ggml-hexagon/htp/htp-tensor.h) flushes only modified tensor address ranges,
- ensuring peer devices and the host CPU observe consistent data in DDR.
+ - [`htp_tensor_flush_all()`](../../../ggml/src/ggml-hexagon/htp/htp-tensor.h) flushes modified tensor address ranges, or the
+ full D-cache when their total size exceeds the flush threshold, ensuring peer devices and the host CPU observe consistent
+ data in DDR.
- Never signal completion before all DMA transfers are drained and dirty tensor flushes have completed.
-
diff --git a/ggml/src/ggml-hexagon/ggml-hexagon.cpp b/ggml/src/ggml-hexagon/ggml-hexagon.cpp
index 1e15e2eb7..8c650b85c 100644
--- a/ggml/src/ggml-hexagon/ggml-hexagon.cpp
+++ b/ggml/src/ggml-hexagon/ggml-hexagon.cpp
@@ -71,6 +71,8 @@
#include "htp/ssm-conv.h"
#include "htp/gated-delta-net-ops.h"
#include "htp/argsort-ops.h"
+#include "htp/concat-ops.h"
+#include "htp/cpy-ops.h"
#include "htp_iface.h"
#include "htp-drv.h"
@@ -400,6 +402,18 @@ static void ggml_hexagon_precompute_pool_2d_params(
bool is_pool_1d
);
+static bool ggml_hexagon_precompute_concat_params(
+ const struct ggml_hexagon_session * sess,
+ const struct ggml_tensor * op,
+ struct htp_concat_kernel_params * kparams
+);
+
+static bool ggml_hexagon_precompute_cpy_params(
+ const struct ggml_hexagon_session * sess,
+ const struct ggml_tensor * op,
+ struct htp_copy_kernel_params * kparams
+);
+
static void ggml_hexagon_precompute_fused_mmnx_params(
const struct ggml_hexagon_session * sess,
const struct ggml_tensor * src0,
@@ -4260,6 +4274,11 @@ void ggml_hexagon_session::enqueue_cpy(const ggml_tensor * src, ggml_tensor * ds
if (with_fence) {
cpy_node.name = "CPY+FENCE";
}
+ const bool ok = ggml_hexagon_precompute_cpy_params(this, node, (struct htp_copy_kernel_params *) cpy_node.kernel_params);
+ const auto * kparams = (const struct htp_copy_kernel_params *) cpy_node.kernel_params;
+ if (ok && !with_fence && kparams->total_elems == 0) {
+ return;
+ }
this->enqueue_op(cpy_node);
}
@@ -6350,6 +6369,208 @@ static void ggml_hexagon_precompute_pool_2d_params(
kparams->inv_kernel_area = 1.0f / (float) (kparams->kernel_x * kparams->kernel_y);
}
+static bool ggml_hexagon_precompute_concat_params(
+ const struct ggml_hexagon_session * sess,
+ const struct ggml_tensor * op,
+ struct htp_concat_kernel_params * kparams
+) {
+ memset(kparams, 0, sizeof(*kparams));
+ kparams->kernel_type = HTP_CONCAT_KERNEL_UNSUPPORTED;
+
+ const struct ggml_tensor * src0 = op->src[0];
+ const struct ggml_tensor * src1 = op->src[1];
+ const struct ggml_tensor * dst = op;
+
+ if (!src0 || !src1 || !dst) {
+ return false;
+ }
+
+ int dim = ((const int32_t *) op->op_params)[0];
+ if (dim < 0 || dim >= GGML_MAX_DIMS) {
+ return false;
+ }
+ kparams->dim = dim;
+
+ if (dst->type != GGML_TYPE_F32 && dst->type != GGML_TYPE_F16 && dst->type != GGML_TYPE_I32) {
+ return false;
+ }
+ if (src0->type != dst->type || src1->type != dst->type) {
+ return false;
+ }
+
+ const uint32_t type_size = ggml_type_size(dst->type);
+
+ for (int d = 0; d < GGML_MAX_DIMS; d++) {
+ const int64_t ne_d = (d == dim) ? src0->ne[d] + src1->ne[d] : src0->ne[d];
+ if (dst->ne[d] != ne_d || (d != dim && src1->ne[d] != dst->ne[d])) {
+ return false;
+ }
+ }
+
+ const bool dma_strides_ok = (src0->nb[0] == type_size && src1->nb[0] == type_size && dst->nb[0] == type_size);
+
+ if (dma_strides_ok) {
+ kparams->kernel_type = HTP_CONCAT_KERNEL_REGULAR;
+ kparams->n_threads = 1;
+ return true;
+ }
+
+ const bool is_src1_transposed = (src1->nb[0] > src1->nb[1]);
+ const bool is_src0_transposed = (src0->nb[0] > src0->nb[1]);
+ const bool transposed_rows_ok = (src0->nb[0] == type_size && src1->nb[1] == type_size && dst->nb[0] == type_size);
+
+ if (dim == 0 && is_src1_transposed && !is_src0_transposed && transposed_rows_ok &&
+ (dst->type == GGML_TYPE_F32 || dst->type == GGML_TYPE_F16)) {
+
+ const uint32_t n_threads = sess->n_threads > 0 ? (uint32_t) sess->n_threads : 8;
+ struct htp_concat_transposed_vtcm_layout layout;
+ htp_concat_transposed_vtcm_layout_build(&layout, (uint32_t) src0->ne[0], (uint32_t) src1->ne[0], type_size, n_threads);
+
+ if (sess->vtcm_size > 0 && layout.total_bytes > sess->vtcm_size) {
+ return false;
+ }
+
+ kparams->kernel_type = HTP_CONCAT_KERNEL_TRANSPOSED;
+ kparams->n_threads = n_threads;
+ kparams->vtcm_size = layout.total_bytes;
+ kparams->spad0_size_per_thread = layout.src0_spad_size_per_thread;
+ kparams->spad1_size_per_thread = layout.src1_spad_size_per_thread;
+ return true;
+ }
+
+ return false;
+}
+
+static bool ggml_hexagon_precompute_cpy_params(
+ const struct ggml_hexagon_session * sess,
+ const struct ggml_tensor * op,
+ struct htp_copy_kernel_params * kparams
+) {
+ memset(kparams, 0, sizeof(*kparams));
+ kparams->kernel_type = HTP_COPY_KERNEL_UNSUPPORTED;
+
+ const struct ggml_tensor * src0 = op->src[0];
+ const struct ggml_tensor * dst = op;
+
+ if (!src0 || !dst) {
+ return false;
+ }
+
+ if (src0->type != GGML_TYPE_F32 && src0->type != GGML_TYPE_F16 && src0->type != GGML_TYPE_I32) {
+ return false;
+ }
+ if (dst->type != GGML_TYPE_F32 && dst->type != GGML_TYPE_F16 && dst->type != GGML_TYPE_I32) {
+ return false;
+ }
+
+ const int64_t nelem_src = ggml_nelements(src0);
+ const int64_t nelem_dst = ggml_nelements(dst);
+ if (nelem_src != nelem_dst || nelem_src < 0) {
+ return false;
+ }
+
+ const uint32_t src_type_size = ggml_type_size(src0->type);
+ const uint32_t dst_type_size = ggml_type_size(dst->type);
+
+ kparams->src0_type_size = (uint8_t) src_type_size;
+ kparams->dst_type_size = (uint8_t) dst_type_size;
+ kparams->total_elems = (uint32_t) nelem_src;
+
+ if (nelem_src == 0) {
+ kparams->kernel_type = HTP_COPY_KERNEL_1D_CONTIG;
+ return true;
+ }
+
+ if (nelem_src == 1) {
+ if (src0->type == dst->type) {
+ kparams->kernel_type = HTP_COPY_KERNEL_SCALAR;
+ return true;
+ }
+ if ((src0->type == GGML_TYPE_F32 && dst->type == GGML_TYPE_I32) ||
+ (src0->type == GGML_TYPE_I32 && dst->type == GGML_TYPE_F32) ||
+ (src0->type == GGML_TYPE_F32 && dst->type == GGML_TYPE_F16) ||
+ (src0->type == GGML_TYPE_F16 && dst->type == GGML_TYPE_F32)) {
+ kparams->kernel_type = HTP_COPY_KERNEL_SCALAR;
+ return true;
+ }
+ return false;
+ }
+
+ const bool sametype = (src0->type == dst->type);
+
+ bool same_extents = true;
+ for (int d = 0; d < GGML_MAX_DIMS; d++) {
+ if (src0->ne[d] != dst->ne[d]) {
+ same_extents = false;
+ break;
+ }
+ }
+
+ const bool transposed = (src0->nb[0] > src0->nb[1]) || (dst->nb[0] > dst->nb[1]) ||
+ (src0->nb[0] != src_type_size) || (dst->nb[0] != dst_type_size) ||
+ (src0->nb[1] < (size_t) src0->ne[0] * src_type_size) || (dst->nb[1] < (size_t) dst->ne[0] * dst_type_size);
+ const bool sameshape = same_extents && !transposed;
+
+ const bool src_is_contiguous = ggml_is_contiguous(src0);
+ const bool dst_is_contiguous = ggml_is_contiguous(dst);
+
+ if (sametype) {
+ if (src_is_contiguous && dst_is_contiguous) {
+ kparams->kernel_type = HTP_COPY_KERNEL_1D_CONTIG;
+ return true;
+ }
+
+ if (sameshape) {
+ kparams->kernel_type = HTP_COPY_KERNEL_SAMESHAPE_SAMETYPE;
+ kparams->total_rows = (uint32_t) (src0->ne[1] * src0->ne[2] * src0->ne[3]);
+ return true;
+ }
+
+ kparams->kernel_type = HTP_COPY_KERNEL_RESHAPE;
+ kparams->n_threads = sess->n_threads > 0 ? (uint8_t) sess->n_threads : 4;
+ kparams->u.reshape.div_ne0 = init_fastdiv_values((uint32_t) dst->ne[0]);
+ kparams->u.reshape.div_ne1_ne0 = init_fastdiv_values((uint32_t) (dst->ne[1] * dst->ne[0]));
+ kparams->u.reshape.div_ne2_ne1_ne0 = init_fastdiv_values((uint32_t) (dst->ne[2] * dst->ne[1] * dst->ne[0]));
+ kparams->u.reshape.div_ne00 = init_fastdiv_values((uint32_t) src0->ne[0]);
+ kparams->u.reshape.div_ne01_ne00 = init_fastdiv_values((uint32_t) (src0->ne[1] * src0->ne[0]));
+ kparams->u.reshape.div_ne02_ne01_ne00 = init_fastdiv_values((uint32_t) (src0->ne[2] * src0->ne[1] * src0->ne[0]));
+ return true;
+ }
+
+ if (!sameshape) {
+ return false;
+ }
+
+ const bool valid_conversion = (src0->type == GGML_TYPE_F32 && dst->type == GGML_TYPE_F16) ||
+ (src0->type == GGML_TYPE_F16 && dst->type == GGML_TYPE_F32) ||
+ (src0->type == GGML_TYPE_F32 && dst->type == GGML_TYPE_I32) ||
+ (src0->type == GGML_TYPE_I32 && dst->type == GGML_TYPE_F32);
+ if (!valid_conversion) {
+ return false;
+ }
+
+ const uint32_t n_threads = sess->n_threads > 0 ? (uint32_t) sess->n_threads : 4;
+ struct htp_copy_convert_vtcm_layout layout;
+ htp_copy_convert_vtcm_layout_build(&layout, (uint32_t) src0->ne[0], src_type_size, dst_type_size, n_threads);
+
+ if (sess->vtcm_size > 0 && layout.total_bytes > sess->vtcm_size) {
+ return false;
+ }
+
+ kparams->kernel_type = HTP_COPY_KERNEL_SAMESHAPE_CONVERT;
+ kparams->total_rows = (uint32_t) (src0->ne[1] * src0->ne[2] * src0->ne[3]);
+ kparams->n_threads = (uint8_t) n_threads;
+ kparams->vtcm_size = layout.total_bytes;
+ kparams->u.convert.src0_buf_size = layout.src0_buf_size;
+ kparams->u.convert.dst_buf_size = layout.dst_buf_size;
+ kparams->u.convert.spad0_size_per_thread = layout.spad0_size_per_thread;
+ kparams->u.convert.spad1_size_per_thread = layout.spad1_size_per_thread;
+ kparams->u.convert.div_ne01 = init_fastdiv_values((uint32_t) src0->ne[1]);
+ kparams->u.convert.div_ne02_ne01 = init_fastdiv_values((uint32_t) (src0->ne[2] * src0->ne[1]));
+
+ return true;
+}
+
static void ggml_hexagon_precompute_fused_mmnx_params(
const struct ggml_hexagon_session * sess,
const struct ggml_tensor * src0, // W0
@@ -7416,6 +7637,7 @@ static htp_op_code op_remap_to_htp(const ggml_tensor * t) {
case GGML_OP_ADD_ID: return HTP_OP_ADD_ID;
case GGML_OP_SUB: return HTP_OP_SUB;
case GGML_OP_DIV: return HTP_OP_DIV;
+ case GGML_OP_DUP:
case GGML_OP_CPY: return HTP_OP_CPY;
case GGML_OP_CONT: return HTP_OP_CPY;
case GGML_OP_GET_ROWS: return HTP_OP_GET_ROWS;
@@ -7716,8 +7938,21 @@ static ggml_status ggml_backend_hexagon_graph_compute(ggml_backend_t backend, gg
ggml_hexagon_precompute_pool_2d_params(
sess, node.node->src[0], node.dst(),
(struct htp_pool_2d_kernel_params *)node.kernel_params,
- node.opcode == HTP_OP_POOL_1D
+ node.opcode == HTP_OP_POOL_1D);
+ } else if (node.opcode == HTP_OP_CONCAT) {
+ ggml_hexagon_precompute_concat_params(sess,
+ node.node,
+ (struct htp_concat_kernel_params *) node.kernel_params
);
+ } else if (node.opcode == HTP_OP_CPY || node.opcode == HTP_OP_CPY_FENCE) {
+ const bool ok = ggml_hexagon_precompute_cpy_params(sess,
+ node.node,
+ (struct htp_copy_kernel_params *) node.kernel_params
+ );
+ const auto * kparams = (const struct htp_copy_kernel_params *) node.kernel_params;
+ if (ok && node.opcode == HTP_OP_CPY && kparams->total_elems == 0) {
+ continue;
+ }
}
computed_nodes.push_back(std::move(node));
}
@@ -8346,49 +8581,13 @@ static ggml_backend_buffer_type_t ggml_backend_hexagon_device_get_host_buffer_ty
}
static bool ggml_hexagon_supported_cpy(const struct ggml_hexagon_session * sess, const struct ggml_tensor * op) {
- GGML_UNUSED(sess);
-
- const struct ggml_tensor * src0 = op->src[0];
- const struct ggml_tensor * dst = op;
-
- if (src0->type != GGML_TYPE_F32 && src0->type != GGML_TYPE_F16 &&
- src0->type != GGML_TYPE_I32) return false;
- if (dst->type != GGML_TYPE_F32 && dst->type != GGML_TYPE_F16 &&
- dst->type != GGML_TYPE_I32) return false;
-
- const bool is_scalar = (ggml_nelements(src0) == 1 && ggml_nelements(dst) == 1);
- const bool sametype = (src0->type == dst->type);
- const bool transposed = !is_scalar && (ggml_is_transposed(src0) || ggml_is_transposed(dst));
- const bool sameshape = is_scalar || (!transposed && ggml_are_same_shape(src0, dst));
-
- if (src0->type == GGML_TYPE_I32 || dst->type == GGML_TYPE_I32) {
- if (!sameshape) return false;
- if (sametype) return true;
- if ((src0->type == GGML_TYPE_F32 && dst->type == GGML_TYPE_I32) ||
- (src0->type == GGML_TYPE_I32 && dst->type == GGML_TYPE_F32)) {
- return true;
- }
- return false;
- }
-
- // can handle any shape and any same-type (pretty slow if reshaping is required)
- if (sametype) return true;
-
- // cannot handle re-shaping and type conversion at the same time
- if (!sameshape) return false;
-
- return true;
+ struct htp_copy_kernel_params kparams;
+ return ggml_hexagon_precompute_cpy_params(sess, op, &kparams);
}
static bool ggml_hexagon_supported_cont(const struct ggml_hexagon_session * sess, const struct ggml_tensor * op) {
- GGML_UNUSED(sess);
- const struct ggml_tensor * src0 = op->src[0];
-
- // CONT is same-type only and supports F32, F16, and I32.
- if (src0->type != GGML_TYPE_F32 && src0->type != GGML_TYPE_F16 &&
- src0->type != GGML_TYPE_I32) return false;
-
- return true;
+ struct htp_copy_kernel_params kparams;
+ return ggml_hexagon_precompute_cpy_params(sess, op, &kparams);
}
static bool ggml_hexagon_supported_repeat(const struct ggml_hexagon_session * sess, const struct ggml_tensor * op) {
@@ -8415,23 +8614,8 @@ static bool ggml_hexagon_supported_repeat(const struct ggml_hexagon_session * se
}
static bool ggml_hexagon_supported_concat(const struct ggml_hexagon_session * sess, const struct ggml_tensor * op) {
- int dim = ((const int32_t *) op->op_params)[0];
- if (dim < 0 || dim >= GGML_MAX_DIMS) {
- return false;
- }
-
- for (int i = 0; i < GGML_MAX_SRC; ++i) {
- const struct ggml_tensor * src = op->src[i];
- if (!src) {
- continue;
- }
- if (src->type != GGML_TYPE_F32 && src->type != GGML_TYPE_I32 && src->type != GGML_TYPE_F16) {
- return false;
- }
- }
-
- return true;
- GGML_UNUSED(sess);
+ struct htp_concat_kernel_params kparams;
+ return ggml_hexagon_precompute_concat_params(sess, op, &kparams);
}
static bool ggml_hexagon_supported_fill(const struct ggml_hexagon_session * sess, const struct ggml_tensor * op) {
@@ -8601,6 +8785,7 @@ static bool ggml_backend_hexagon_device_supports_op(ggml_backend_dev_t dev, cons
supp = ggml_hexagon_supported_get_rows(sess, op);
break;
+ case GGML_OP_DUP:
case GGML_OP_CPY:
supp = ggml_hexagon_supported_cpy(sess, op);
break;
diff --git a/ggml/src/ggml-hexagon/htp/concat-ops.c b/ggml/src/ggml-hexagon/htp/concat-ops.c
index 4dc146393..f7dda1408 100644
--- a/ggml/src/ggml-hexagon/htp/concat-ops.c
+++ b/ggml/src/ggml-hexagon/htp/concat-ops.c
@@ -1,11 +1,13 @@
+#include "concat-ops.h"
#include "dma-queue.h"
#include "hex-common.h"
-#include "hex-cpy-dma.h"
+#include "dma-copy.h"
#include "hex-fastdiv.h"
#include "hex-profile.h"
#include "hexagon_protos.h"
#include "hexagon_types.h"
#include "htp-ctx.h"
+#include "htp-fence.h"
#include "htp-ops.h"
#include "htp-tensor.h"
#include "htp-vtcm.h"
@@ -16,15 +18,14 @@
struct htp_concat_context {
struct htp_ops_context * octx;
- uint32_t dim;
- uint32_t nrows_per_thread;
- uint32_t row_start;
- uint32_t nrows;
- uint32_t elem_start;
- uint32_t nelems;
- uint32_t nplanes;
- struct fastdiv_values div_ne0;
- struct fastdiv_values div_ne1;
+ uint8_t * spad0_base;
+ uint8_t * spad1_base;
+ uint32_t spad0_size_per_thread;
+ uint32_t spad1_size_per_thread;
+ uint32_t row_start;
+ uint32_t nrows;
+ uint32_t nrows_per_thread;
+ uint32_t nplanes;
struct fastdiv_values div_ne2;
};
@@ -52,8 +53,8 @@ static void concat_2d_f32_transposed(unsigned int nth, unsigned int ith, void *
dma_queue * dma_q = octx->ctx->dma[ith];
- uint8_t * spad0_base = octx->src0_spad.data + ith * octx->src0_spad.size_per_thread;
- uint8_t * spad1_base = octx->src1_spad.data + ith * octx->src1_spad.size_per_thread;
+ uint8_t * spad0_base = cctx->spad0_base + ith * cctx->spad0_size_per_thread;
+ uint8_t * spad1_base = cctx->spad1_base + ith * cctx->spad1_size_per_thread;
const uint32_t block_i = 32;
const uint32_t spad1_stride = block_i * sizeof(float);
@@ -127,6 +128,7 @@ static void concat_2d_f32_transposed(unsigned int nth, unsigned int ith, void *
p = np;
i = ni;
}
+ dma_queue_flush(dma_q);
}
static void concat_2d_f16_transposed(unsigned int nth, unsigned int ith, void * data) {
@@ -147,8 +149,8 @@ static void concat_2d_f16_transposed(unsigned int nth, unsigned int ith, void *
dma_queue * dma_q = octx->ctx->dma[ith];
- uint8_t * spad0_base = octx->src0_spad.data + ith * octx->src0_spad.size_per_thread;
- uint8_t * spad1_base = octx->src1_spad.data + ith * octx->src1_spad.size_per_thread;
+ uint8_t * spad0_base = cctx->spad0_base + ith * cctx->spad0_size_per_thread;
+ uint8_t * spad1_base = cctx->spad1_base + ith * cctx->spad1_size_per_thread;
const uint32_t block_i = 64;
const uint32_t spad1_stride = block_i * sizeof(__fp16);
@@ -222,219 +224,123 @@ static void concat_2d_f16_transposed(unsigned int nth, unsigned int ith, void *
p = np;
i = ni;
}
+ dma_queue_flush(dma_q);
}
-static void concat_generic(unsigned int nth, unsigned int ith, void * data) {
- struct htp_concat_context * cctx = (struct htp_concat_context *) data;
- struct htp_ops_context * octx = cctx->octx;
-
- const struct htp_tensor * src0 = octx->src[0];
- const struct htp_tensor * src1 = octx->src[1];
- const struct htp_tensor * dst = octx->dst;
-
- const int dim = cctx->dim;
- const uint32_t type_size = (dst->type == HTP_TYPE_F32 || dst->type == HTP_TYPE_I32) ? 4 : 2;
-
- const uint32_t ne[4] = {dst->ne[0], dst->ne[1], dst->ne[2], dst->ne[3]};
-
- // Per-device element range aligned to prevent false sharing
- const uint32_t elem_start = cctx->elem_start;
- const uint32_t nelems = cctx->nelems;
- const uint32_t chunk_size = fastdiv(nelems + nth - 1, &octx->n_threads_div);
-
- const uint32_t start_idx = MIN(elem_start + ith * chunk_size, elem_start + nelems);
- const uint32_t end_idx = MIN(start_idx + chunk_size, elem_start + nelems);
-
- // Naive scalar element-wise copy
- for (uint32_t idx = start_idx; idx < end_idx; idx++) {
- uint32_t idx_div_ne0 = fastdiv(idx, &cctx->div_ne0);
- uint32_t i0 = idx - idx_div_ne0 * ne[0];
-
- uint32_t idx_div_ne01 = fastdiv(idx_div_ne0, &cctx->div_ne1);
- uint32_t i1 = idx_div_ne0 - idx_div_ne01 * ne[1];
-
- uint32_t idx_div_ne012 = fastdiv(idx_div_ne01, &cctx->div_ne2);
- uint32_t i2 = idx_div_ne01 - idx_div_ne012 * ne[2];
- uint32_t i3 = idx_div_ne012;
-
- uint8_t * dst_ptr = (uint8_t *)dst->data + i3 * dst->nb[3] + i2 * dst->nb[2] + i1 * dst->nb[1] + i0 * dst->nb[0];
-
- uint32_t idx_dim = 0;
- if (dim == 0) idx_dim = i0;
- else if (dim == 1) idx_dim = i1;
- else if (dim == 2) idx_dim = i2;
- else if (dim == 3) idx_dim = i3;
-
- const struct htp_tensor * src = (idx_dim < src0->ne[dim]) ? src0 : src1;
-
- uint32_t s0 = i0;
- uint32_t s1 = i1;
- uint32_t s2 = i2;
- uint32_t s3 = i3;
-
- if (dim == 0 && src == src1) s0 -= src0->ne[0];
- if (dim == 1 && src == src1) s1 -= src0->ne[1];
- if (dim == 2 && src == src1) s2 -= src0->ne[2];
- if (dim == 3 && src == src1) s3 -= src0->ne[3];
-
- uint8_t * src_ptr = (uint8_t *)src->data + s3 * src->nb[3] + s2 * src->nb[2] + s1 * src->nb[1] + s0 * src->nb[0];
-
- if (type_size == 4) {
- *(float*)dst_ptr = *(float*)src_ptr;
- } else {
- *(__fp16*)dst_ptr = *(__fp16*)src_ptr;
- }
- }
-}
-
-static bool concat_dma(struct htp_ops_context * octx, int dim, uint32_t type_size) {
- if (dim < 0 || dim >= HTP_OP_MAX_DIMS) {
- return false;
- }
-
+static int concat_regular(struct htp_ops_context * octx, int dim, uint32_t type_size) {
const struct htp_tensor * src0 = octx->src[0];
const struct htp_tensor * src1 = octx->src[1];
const struct htp_tensor * dst = octx->dst;
- // Not partitioned across devices: the row/element-split paths handle that.
- if (octx->ctx->mdev.count > 1 ||
- (dst->type != HTP_TYPE_F32 && dst->type != HTP_TYPE_F16 && dst->type != HTP_TYPE_I32) ||
- src0->type != dst->type || src1->type != dst->type || src0->nb[0] != type_size || src1->nb[0] != type_size ||
- dst->nb[0] != type_size || (size_t) dst->ne[0] * type_size > DMA_MAX_SIZE_24B ||
- dst->nb[1] > DMA_MAX_STRIDE_24B || src0->nb[1] > DMA_MAX_STRIDE_24B || src1->nb[1] > DMA_MAX_STRIDE_24B) {
- return false;
- }
-
- for (int d = 0; d < HTP_OP_MAX_DIMS; d++) {
- const uint32_t ne_d = (d == dim) ? src0->ne[d] + src1->ne[d] : src0->ne[d];
- if (dst->ne[d] != ne_d || (d != dim && src1->ne[d] != dst->ne[d])) {
- return false;
- }
- }
-
- // The two views of dst, shaped like the sources.
struct htp_tensor view0 = *dst;
struct htp_tensor view1 = *dst;
for (int d = 0; d < HTP_OP_MAX_DIMS; d++) {
view0.ne[d] = src0->ne[d];
view1.ne[d] = src1->ne[d];
}
- view1.data += (uint64_t) src0->ne[dim] * dst->nb[dim];
+ view1.data += src0->ne[dim] * dst->nb[dim];
+
+ const uint32_t total_rows_0 = src0->ne[1] * src0->ne[2] * src0->ne[3];
+ const uint32_t total_rows_1 = src1->ne[1] * src1->ne[2] * src1->ne[3];
+
+ uint32_t rstart0 = 0, nrows0 = total_rows_0;
+ uint32_t rstart1 = 0, nrows1 = total_rows_1;
+
+ if (octx->ctx->mdev.count > 1) {
+ const struct htp_tensor_mdev_range range0 = htp_tensor_mdev_partition(
+ total_rows_0, 1, octx->ctx->mdev.idx, octx->ctx->mdev.count, &octx->ctx->mdev.count_div);
+ rstart0 = range0.start;
+ nrows0 = range0.count;
+
+ const struct htp_tensor_mdev_range range1 = htp_tensor_mdev_partition(
+ total_rows_1, 1, octx->ctx->mdev.idx, octx->ctx->mdev.count, &octx->ctx->mdev.count_div);
+ rstart1 = range1.start;
+ nrows1 = range1.count;
+ }
dma_queue * q = octx->ctx->dma[0];
- cpy_dma_sametype_sameshape(q, &view0, src0, type_size);
- cpy_dma_sametype_sameshape(q, &view1, src1, type_size);
+ dma_cpy_sametype_sameshape_range(q, &view0, src0, type_size, rstart0, nrows0);
+ dma_cpy_sametype_sameshape_range(q, &view1, src1, type_size, rstart1, nrows1);
dma_queue_flush(q);
- return true;
+
+ return HTP_STATUS_OK;
}
-int op_concat(struct htp_ops_context * octx) {
- int dim = octx->op_params[0];
- if (dim < 0 || dim >= HTP_OP_MAX_DIMS) {
- return HTP_STATUS_NO_SUPPORT;
+static int concat_transposed(struct htp_ops_context * octx, const struct htp_concat_kernel_params * kparams, uint32_t type_size) {
+ if (!htp_ops_context_set_n_threads(octx, kparams->n_threads)) {
+ return HTP_STATUS_INVAL_PARAMS;
}
- const struct htp_tensor * src0 = octx->src[0];
- const struct htp_tensor * src1 = octx->src[1];
- const struct htp_tensor * dst = octx->dst;
+ const struct htp_tensor * dst = octx->dst;
- const uint32_t type_size = (dst->type == HTP_TYPE_F32 || dst->type == HTP_TYPE_I32) ? 4 : 2;
- bool is_src1_transposed = (src1->nb[0] > src1->nb[1]);
- bool is_src0_transposed = (src0->nb[0] > src0->nb[1]);
+ const uint32_t total_rows = dst->ne[1];
+ uint32_t row_start = 0;
+ uint32_t nrows = total_rows;
+ if (octx->ctx->mdev.count > 1) {
+ const struct htp_tensor_mdev_range range = htp_tensor_mdev_partition(total_rows, 1, octx->ctx->mdev.idx, octx->ctx->mdev.count, &octx->ctx->mdev.count_div);
+ row_start = range.start;
+ nrows = range.count;
+ }
- if (concat_dma(octx, dim, type_size)) {
+ if (nrows == 0 || dst->ne[2] == 0 || dst->ne[3] == 0) {
return HTP_STATUS_OK;
}
- uint32_t n_threads = octx->n_threads;
- struct htp_concat_context cctx;
- cctx.octx = octx;
- cctx.dim = dim;
- cctx.div_ne0 = init_fastdiv_values(dst->ne[0]);
- cctx.div_ne1 = init_fastdiv_values(dst->ne[1]);
- cctx.div_ne2 = init_fastdiv_values(dst->ne[2]);
-
- void (*worker_func)(unsigned int, unsigned int, void *) = concat_generic;
-
- const bool rows_ok = src0->nb[0] == type_size && src1->nb[1] == type_size && dst->nb[0] == type_size;
-
- if (dim == 0 && is_src1_transposed && !is_src0_transposed && rows_ok) {
- const uint32_t total_rows = dst->ne[1];
- const size_t dst_data_row_size = dst->ne[0] * type_size;
- uint32_t row_start = 0;
- uint32_t nrows = total_rows;
- if (octx->ctx->mdev.count > 1) {
- uint32_t rows_per_chunk = 0;
- htp_tensor_mdev_rows_per_chunk(dst, type_size, (uint32_t) dst_data_row_size, &rows_per_chunk);
- const struct htp_tensor_mdev_range range = htp_tensor_mdev_partition(total_rows, rows_per_chunk, octx->ctx->mdev.idx, octx->ctx->mdev.count, &octx->ctx->mdev.count_div);
- row_start = range.start;
- nrows = range.count;
- }
-
- if (nrows == 0) {
- return HTP_STATUS_OK;
- }
-
- cctx.row_start = row_start;
- cctx.nrows = nrows;
- cctx.nplanes = dst->ne[2] * dst->ne[3];
-
- uint32_t block_i = (type_size == 4) ? 32 : 64;
-
- cctx.nrows_per_thread = fastdiv(nrows + n_threads - 1, &octx->n_threads_div);
+ if (kparams->vtcm_size > octx->ctx->vtcm_size) {
+ return HTP_STATUS_VTCM_TOO_SMALL;
+ }
- // Allocate VTCM
- uint32_t spad1_stride = block_i * type_size;
+ const uint32_t n_threads = octx->n_threads;
- uint32_t src1_ne0_padded = hex_round_up(src1->ne[0], block_i);
- // src0 row is right-aligned to VLEN so the gathered src1 part starts aligned
- uint32_t spad0_row_bytes = hex_round_up(src0->ne[0] * type_size, VLEN) + src1_ne0_padded * type_size;
+ // layout precomputed on host; kept for reference:
+ // struct htp_concat_transposed_vtcm_layout layout;
+ // htp_concat_transposed_vtcm_layout_build(&layout, octx->src[0]->ne[0], octx->src[1]->ne[0], type_size, n_threads);
- octx->src0_spad.size_per_thread = block_i * spad0_row_bytes;
- octx->src1_spad.size_per_thread = src1_ne0_padded * spad1_stride;
+ uint8_t * vtcm_base = (uint8_t *) octx->ctx->vtcm_base;
- octx->src0_spad.size = n_threads * octx->src0_spad.size_per_thread;
- octx->src1_spad.size = n_threads * octx->src1_spad.size_per_thread;
+ struct htp_concat_context cctx;
+ cctx.octx = octx;
+ cctx.spad0_base = vtcm_base;
+ cctx.spad1_base = vtcm_base + n_threads * kparams->spad0_size_per_thread;
+ cctx.spad0_size_per_thread = kparams->spad0_size_per_thread;
+ cctx.spad1_size_per_thread = kparams->spad1_size_per_thread;
+ cctx.row_start = row_start;
+ cctx.nrows = nrows;
+ cctx.nplanes = dst->ne[2] * dst->ne[3];
+ cctx.div_ne2 = init_fastdiv_values(dst->ne[2]);
+ cctx.nrows_per_thread = fastdiv(nrows + n_threads - 1, &octx->n_threads_div);
+
+ work_queue_func_t worker_func = (type_size == 4) ? concat_2d_f32_transposed : concat_2d_f16_transposed;
+ work_queue_run(octx->ctx->work_queue, worker_func, &cctx, n_threads);
+ return HTP_STATUS_OK;
+}
- if (octx->src0_spad.size + octx->src1_spad.size > octx->ctx->vtcm_size) {
- return HTP_STATUS_VTCM_TOO_SMALL;
- }
+int op_concat(struct htp_ops_context * octx) {
+ const struct htp_concat_kernel_params * kparams = (const struct htp_concat_kernel_params *) octx->kernel_params;
+ const struct htp_tensor * dst = octx->dst;
+ const uint32_t type_size = (dst->type == HTP_TYPE_F32 || dst->type == HTP_TYPE_I32) ? 4 : 2;
- octx->src0_spad.data = octx->ctx->vtcm_base;
- octx->src1_spad.data = octx->src0_spad.data + octx->src0_spad.size;
- octx->src0_spad.src = NULL;
- octx->src1_spad.src = NULL;
+ int status = HTP_STATUS_OK;
+ switch (kparams->kernel_type) {
+ case HTP_CONCAT_KERNEL_REGULAR:
+ status = concat_regular(octx, kparams->dim, type_size);
+ break;
- if (type_size == 4) {
- worker_func = concat_2d_f32_transposed;
- } else {
- worker_func = concat_2d_f16_transposed;
- }
- } else {
- if (htp_tensor_is_extended(src0) || htp_tensor_is_extended(src1) || htp_tensor_is_extended(dst)) {
- return HTP_STATUS_NO_SUPPORT;
- }
+ case HTP_CONCAT_KERNEL_TRANSPOSED:
+ status = concat_transposed(octx, kparams, type_size);
+ break;
- const uint32_t total_elements = dst->ne[0] * dst->ne[1] * dst->ne[2] * dst->ne[3];
- uint32_t elem_start = 0;
- uint32_t nelems = total_elements;
- if (octx->ctx->mdev.count > 1) {
- const uint32_t elems_per_chunk = HEX_L2_LINE_SIZE / type_size;
- const bool can_split = htp_tensor_mdev_data_aligned(dst) && htp_tensor_is_contiguous(dst, type_size) && !htp_tensor_is_permuted(dst);
- const struct htp_tensor_mdev_range range = htp_tensor_mdev_partition(total_elements, can_split ? elems_per_chunk : 0, octx->ctx->mdev.idx, octx->ctx->mdev.count, &octx->ctx->mdev.count_div);
- elem_start = range.start;
- nelems = range.count;
- }
+ default:
+ status = HTP_STATUS_NO_SUPPORT;
+ break;
+ }
- if (nelems == 0) {
- return HTP_STATUS_OK;
- }
+ htp_ops_context_set_status(octx, status);
- cctx.elem_start = elem_start;
- cctx.nelems = nelems;
+ if (octx->ctx->mdev.count > 1) {
+ htp_mdev_group_barrier(octx);
}
- work_queue_run(octx->ctx->work_queue, worker_func, &cctx, n_threads);
- return HTP_STATUS_OK;
+ return octx->status;
}
diff --git a/ggml/src/ggml-hexagon/htp/concat-ops.h b/ggml/src/ggml-hexagon/htp/concat-ops.h
new file mode 100644
index 000000000..1d7218441
--- /dev/null
+++ b/ggml/src/ggml-hexagon/htp/concat-ops.h
@@ -0,0 +1,53 @@
+#ifndef HTP_CONCAT_OPS_H
+#define HTP_CONCAT_OPS_H
+
+#include "hex-common.h"
+#include <stdint.h>
+
+enum htp_concat_kernel_type {
+ HTP_CONCAT_KERNEL_UNSUPPORTED = 0,
+ HTP_CONCAT_KERNEL_REGULAR = 1,
+ HTP_CONCAT_KERNEL_TRANSPOSED = 2,
+};
+
+struct htp_concat_kernel_params {
+ uint8_t kernel_type;
+ uint8_t dim;
+ uint8_t n_threads;
+ uint8_t pad;
+
+ uint32_t vtcm_size;
+ uint32_t spad0_size_per_thread;
+ uint32_t spad1_size_per_thread;
+};
+
+#if defined(__cplusplus)
+static_assert(sizeof(struct htp_concat_kernel_params) <= 128, "htp_concat_kernel_params is too large for kernel_params blob");
+#else
+_Static_assert(sizeof(struct htp_concat_kernel_params) <= 128, "htp_concat_kernel_params is too large for kernel_params blob");
+#endif
+
+struct htp_concat_transposed_vtcm_layout {
+ uint32_t src0_spad_size_per_thread;
+ uint32_t src1_spad_size_per_thread;
+ uint32_t total_bytes;
+};
+
+static inline void htp_concat_transposed_vtcm_layout_build(
+ struct htp_concat_transposed_vtcm_layout * layout,
+ uint32_t src0_ne0,
+ uint32_t src1_ne0,
+ uint32_t type_size,
+ uint32_t n_threads) {
+
+ uint32_t block_i = (type_size == 4) ? 32 : 64;
+ uint32_t spad1_stride = block_i * type_size;
+ uint32_t src1_ne0_padded = hex_round_up(src1_ne0, block_i);
+ uint32_t spad0_row_bytes = hex_round_up(src0_ne0 * type_size, 128) + src1_ne0_padded * type_size;
+
+ layout->src0_spad_size_per_thread = block_i * spad0_row_bytes;
+ layout->src1_spad_size_per_thread = src1_ne0_padded * spad1_stride;
+ layout->total_bytes = n_threads * (layout->src0_spad_size_per_thread + layout->src1_spad_size_per_thread);
+}
+
+#endif // HTP_CONCAT_OPS_H
diff --git a/ggml/src/ggml-hexagon/htp/cpy-ops.c b/ggml/src/ggml-hexagon/htp/cpy-ops.c
index df38d3eed..d366acb75 100644
--- a/ggml/src/ggml-hexagon/htp/cpy-ops.c
+++ b/ggml/src/ggml-hexagon/htp/cpy-ops.c
@@ -11,7 +11,8 @@
#define GGML_COMMON_DECL_C
#include "ggml-common.h"
-#include "hex-cpy-dma.h"
+#include "cpy-ops.h"
+#include "dma-copy.h"
#include "htp-ctx.h"
#include "htp-fence.h"
#include "htp-ops.h"
@@ -19,34 +20,19 @@
#include "hvx-utils.h"
struct htp_copy_context {
- struct htp_ops_context * octx;
+ struct htp_ops_context * octx;
+ const struct htp_copy_kernel_params * kparams;
- uint32_t src0_type_size;
- uint32_t src0_block_size;
+ uint32_t row_start;
+ uint32_t nrows;
+ uint32_t src0_nrows_per_thread;
- uint32_t dst_type_size;
- uint32_t dst_block_size;
+ uint32_t elem_start;
+ uint32_t nelem;
+ uint32_t elem_per_thread;
- uint32_t src0_blocks_per_row;
- uint32_t dst_blocks_per_row;
-
- uint32_t elem_start;
- uint32_t nelem;
- uint32_t elem_per_thread;
-
- uint32_t src0_nrows_per_thread;
- uint32_t row_start;
- uint32_t nrows;
-
- struct fastdiv_values div_ne01;
- struct fastdiv_values div_ne02_ne01;
-
- struct fastdiv_values div_ne0;
- struct fastdiv_values div_ne1_ne0;
- struct fastdiv_values div_ne2_ne1_ne0;
- struct fastdiv_values div_ne00;
- struct fastdiv_values div_ne01_ne00;
- struct fastdiv_values div_ne02_ne01_ne00;
+ uint8_t * vtcm_src0;
+ uint8_t * vtcm_dst;
};
#define cpy_preamble \
@@ -73,52 +59,6 @@ struct htp_copy_context {
const uint32_t nb2 = dst->nb[2]; \
const uint32_t nb3 = dst->nb[3];
-#define DEFINE_CPY_SAMESHAPE(NAME, ELEM_TYPE, ELEM_SIZE) \
-static void cpy_thread_##NAME##_sameshape(unsigned int nth, unsigned int ith, void * data) { \
- struct htp_copy_context * ct = (struct htp_copy_context *) data; \
- struct htp_ops_context * octx = ct->octx; \
- cpy_preamble; \
- const uint32_t dr = ct->src0_nrows_per_thread; \
- const uint32_t ir0 = ct->row_start + dr * ith; \
- const uint32_t ir1 = MIN(ir0 + dr, ct->row_start + ct->nrows); \
- if (ir0 >= ir1) return; \
- const bool contiguous = htp_tensor_is_contiguous(src0, ELEM_SIZE) && htp_tensor_is_contiguous(dst, ELEM_SIZE); \
- if (contiguous) { \
- dma_queue * dma_q = octx->ctx->dma[ith]; \
- dma_addr_t dst_addr = dst->data + ir0 * ne00 * ELEM_SIZE; \
- dma_addr_t src0_addr = src0->data + ir0 * ne00 * ELEM_SIZE; \
- cpy_dma_sametype_reshape_contig(dma_q, dst_addr, src0_addr, (ir1 - ir0) * ne00 * ELEM_SIZE); \
- dma_queue_flush(dma_q); \
- return; \
- } \
- const uint32_t ne02_ne01 = ne02 * ne01; \
- uint32_t i03 = fastdiv(ir0, &ct->div_ne02_ne01); \
- uint32_t rem = ir0 - i03 * ne02_ne01; \
- uint32_t i02 = fastdiv(rem, &ct->div_ne01); \
- uint32_t i01 = rem - i02 * ne01; \
- uint8_t * dst_ptr = (uint8_t *) dst->data + i01*nb1 + i02*nb2 + i03*nb3; \
- uint8_t * src0_ptr = (uint8_t *) src0->data + i01*nb01 + i02*nb02 + i03*nb03; \
- for (uint32_t r = ir0; r < ir1; r++) { \
- hex_l2fetch(src0_ptr, ne00 * ELEM_SIZE, nb01, 2); \
- hvx_copy_uu(dst_ptr, src0_ptr, ne00, ELEM_SIZE); \
- dst_ptr += nb1; \
- src0_ptr += nb01; \
- if (++i01 == ne01) { \
- i01 = 0; \
- if (++i02 == ne02) { \
- i02 = 0; \
- i03++; \
- } \
- dst_ptr = (uint8_t *) dst->data + i02*nb2 + i03*nb3; \
- src0_ptr = (uint8_t *) src0->data + i02*nb02 + i03*nb03; \
- } \
- } \
-}
-
-DEFINE_CPY_SAMESHAPE(f32, float, 4)
-DEFINE_CPY_SAMESHAPE(f16, __fp16, 2)
-DEFINE_CPY_SAMESHAPE(i32, int32_t, 4)
-
#define DEFINE_CPY_RESHAPE(NAME, ELEM_TYPE, ELEM_SIZE) \
static void cpy_thread_##NAME##_reshape(unsigned int nth, unsigned int ith, void * data) { \
struct htp_copy_context * ct = (struct htp_copy_context *) data; \
@@ -129,52 +69,45 @@ static void cpy_thread_##NAME##_reshape(unsigned int nth, unsigned int ith, void
const uint32_t th_end = MIN(th_start + th_nelem, ct->elem_start + ct->nelem); \
if (th_start >= th_end) return; \
\
- if (htp_tensor_is_contiguous(src0, ELEM_SIZE) && htp_tensor_is_contiguous(dst, ELEM_SIZE)) { \
- dma_queue * dma_q = octx->ctx->dma[ith]; \
- dma_addr_t dst_addr = dst->data + th_start * ELEM_SIZE; \
- dma_addr_t src0_addr = src0->data + th_start * ELEM_SIZE; \
- cpy_dma_sametype_reshape_contig(dma_q, dst_addr, src0_addr, (th_end - th_start) * ELEM_SIZE); \
- dma_queue_flush(dma_q); \
- return; \
- } \
+ dma_queue * dma_q = octx->ctx->dma[ith]; \
\
const uint32_t ne01_ne00 = ne01 * ne00; \
const uint32_t ne02_ne01_ne00 = ne02 * ne01_ne00; \
const uint32_t ne1_ne0 = ne1 * ne0; \
const uint32_t ne2_ne1_ne0 = ne2 * ne1_ne0; \
\
+ const struct htp_copy_reshape_params * rsh = &ct->kparams->u.reshape; \
uint32_t e = th_start; \
- uint32_t i13 = fastdiv(e, &ct->div_ne2_ne1_ne0); \
+ uint32_t i13 = fastdiv(e, &rsh->div_ne2_ne1_ne0); \
uint32_t rem = e - i13 * ne2_ne1_ne0; \
- uint32_t i12 = fastdiv(rem, &ct->div_ne1_ne0); \
+ uint32_t i12 = fastdiv(rem, &rsh->div_ne1_ne0); \
uint32_t rem2 = rem - i12 * ne1_ne0; \
- uint32_t i11 = fastdiv(rem2, &ct->div_ne0); \
+ uint32_t i11 = fastdiv(rem2, &rsh->div_ne0); \
uint32_t i10 = rem2 - i11 * ne0; \
\
- uint32_t i03 = fastdiv(e, &ct->div_ne02_ne01_ne00); \
+ uint32_t i03 = fastdiv(e, &rsh->div_ne02_ne01_ne00); \
uint32_t rem_s = e - i03 * ne02_ne01_ne00; \
- uint32_t i02 = fastdiv(rem_s, &ct->div_ne01_ne00); \
+ uint32_t i02 = fastdiv(rem_s, &rsh->div_ne01_ne00); \
uint32_t rem2_s = rem_s - i02 * ne01_ne00; \
- uint32_t i01 = fastdiv(rem2_s, &ct->div_ne00); \
+ uint32_t i01 = fastdiv(rem2_s, &rsh->div_ne00); \
uint32_t i00 = rem2_s - i01 * ne00; \
\
- char * dst_ptr = (char *) dst->data + i10*nb0 + i11*nb1 + i12*nb2 + i13*nb3; \
- const char * src0_ptr = (const char *) src0->data + i00*nb00 + i01*nb01 + i02*nb02 + i03*nb03; \
+ dma_addr_t dst_addr = dst->data + i10*nb0 + i11*nb1 + i12*nb2 + i13*nb3; \
+ dma_addr_t src0_addr = src0->data + i00*nb00 + i01*nb01 + i02*nb02 + i03*nb03; \
\
const bool rows_contig = (nb00 == ELEM_SIZE) && (nb0 == ELEM_SIZE); \
\
while (e < th_end) { \
- uint32_t run = 1; \
+ const uint32_t run = MIN(MIN(ne00 - i00, ne0 - i10), th_end - e); \
if (rows_contig) { \
- run = MIN(MIN(ne00 - i00, ne0 - i10), th_end - e); \
- hvx_copy_uu((uint8_t *) dst_ptr, (const uint8_t *) src0_ptr, run, ELEM_SIZE); \
+ dma_cpy_sametype_reshape_contig(dma_q, dst_addr, src0_addr, run * ELEM_SIZE); \
} else { \
- *((ELEM_TYPE *) dst_ptr) = *((const ELEM_TYPE *) src0_ptr); \
+ dma_cpy_push_2d_chunked(dma_q, dst_addr, src0_addr, nb0, nb00, ELEM_SIZE, run); \
} \
e += run; \
\
- dst_ptr += run * nb0; \
- i10 += run; \
+ dst_addr += run * nb0; \
+ i10 += run; \
if (i10 == ne0) { \
i10 = 0; \
if (++i11 == ne1) { \
@@ -184,11 +117,11 @@ static void cpy_thread_##NAME##_reshape(unsigned int nth, unsigned int ith, void
i13++; \
} \
} \
- dst_ptr = (char *) dst->data + i11*nb1 + i12*nb2 + i13*nb3; \
+ dst_addr = dst->data + i11*nb1 + i12*nb2 + i13*nb3; \
} \
\
- src0_ptr += run * nb00; \
- i00 += run; \
+ src0_addr += run * nb00; \
+ i00 += run; \
if (i00 == ne00) { \
i00 = 0; \
if (++i01 == ne01) { \
@@ -198,371 +131,350 @@ static void cpy_thread_##NAME##_reshape(unsigned int nth, unsigned int ith, void
i03++; \
} \
} \
- src0_ptr = (const char *) src0->data + i01*nb01 + i02*nb02 + i03*nb03; \
+ src0_addr = src0->data + i01*nb01 + i02*nb02 + i03*nb03; \
} \
} \
+ dma_queue_flush(dma_q); \
}
DEFINE_CPY_RESHAPE(f32, float, 4)
DEFINE_CPY_RESHAPE(f16, __fp16, 2)
DEFINE_CPY_RESHAPE(i32, int32_t, 4)
-static void cpy_thread_f16_f32_sameshape(unsigned int nth, unsigned int ith, void * data) {
- struct htp_copy_context * ct = (struct htp_copy_context *) data;
- struct htp_ops_context * octx = ct->octx;
- cpy_preamble;
-
- const uint32_t dr = ct->src0_nrows_per_thread;
- const uint32_t ir0 = ct->row_start + dr * ith;
- const uint32_t ir1 = MIN(ir0 + dr, ct->row_start + ct->nrows);
- if (ir0 >= ir1) return;
-
- const uint32_t ne02_ne01 = ne02 * ne01;
- uint32_t i03 = fastdiv(ir0, &ct->div_ne02_ne01);
- uint32_t rem = ir0 - i03 * ne02_ne01;
- uint32_t i02 = fastdiv(rem, &ct->div_ne01);
- uint32_t i01 = rem - i02 * ne01;
-
- uint8_t* dst_ptr = (uint8_t*) dst->data + i01*nb1 + i02*nb2 + i03*nb3;
- uint8_t* src0_ptr = (uint8_t*) src0->data + i01*nb01 + i02*nb02 + i03*nb03;
-
- for (uint32_t r = ir0; r < ir1; r++) {
- hex_l2fetch(src0_ptr, ne00 * sizeof(float), nb01, 2);
- hvx_copy_f16_f32_uu(dst_ptr, src0_ptr, ne00);
- dst_ptr += nb1;
- src0_ptr += nb01;
- if (++i01 == ne01) {
- i01 = 0;
- if (++i02 == ne02) {
- i02 = 0;
- i03++;
- }
- dst_ptr = (uint8_t*) dst->data + i02*nb2 + i03*nb3;
- src0_ptr = (uint8_t*) src0->data + i02*nb02 + i03*nb03;
- }
- }
+#define DEFINE_CPY_CONVERT_SAMESHAPE(NAME, CONV_FUNC) \
+static void cpy_thread_##NAME##_sameshape(unsigned int nth, unsigned int ith, void * data) { \
+ struct htp_copy_context * ct = (struct htp_copy_context *) data; \
+ struct htp_ops_context * octx = ct->octx; \
+ cpy_preamble; \
+ \
+ const uint32_t dr = ct->src0_nrows_per_thread; \
+ const uint32_t ir0 = ct->row_start + dr * ith; \
+ const uint32_t ir1 = MIN(ir0 + dr, ct->row_start + ct->nrows); \
+ if (ir0 >= ir1) return; \
+ const uint32_t nrows_thread = ir1 - ir0; \
+ \
+ dma_queue * dma_q = octx->ctx->dma[ith]; \
+ struct htp_thread_trace * tr = &octx->ctx->trace[ith]; \
+ \
+ const struct htp_copy_convert_params * cvt = &ct->kparams->u.convert; \
+ const uint32_t src0_buf_size = cvt->src0_buf_size; \
+ const uint32_t dst_buf_size = cvt->dst_buf_size; \
+ uint8_t * vtcm_src0_base = ct->vtcm_src0 + ith * cvt->spad0_size_per_thread; \
+ uint8_t * vtcm_dst_base = ct->vtcm_dst + ith * cvt->spad1_size_per_thread; \
+ const uint32_t src0_row_size = ne00 * ct->kparams->src0_type_size; \
+ const uint32_t dst_row_size = ne00 * ct->kparams->dst_type_size; \
+ \
+ const uint32_t ne02_ne01 = ne02 * ne01; \
+ uint32_t i03 = fastdiv(ir0, &cvt->div_ne02_ne01); \
+ uint32_t rem = ir0 - i03 * ne02_ne01; \
+ uint32_t i02 = fastdiv(rem, &cvt->div_ne01); \
+ uint32_t i01 = rem - i02 * ne01; \
+ \
+ uint32_t f_i01 = i01, f_i02 = i02, f_i03 = i03; \
+ dma_addr_t f_src0_addr = src0->data + f_i01*nb01 + f_i02*nb02 + f_i03*nb03; \
+ \
+ uint32_t c_i01 = i01, c_i02 = i02, c_i03 = i03; \
+ dma_addr_t c_dst_addr = dst->data + c_i01*nb1 + c_i02*nb2 + c_i03*nb3; \
+ \
+ for (uint32_t r = 0; r < nrows_thread && r < 2; ++r) { \
+ uint8_t * src_spad = vtcm_src0_base + r * src0_buf_size; \
+ uint8_t * dst_spad = vtcm_dst_base + r * dst_buf_size; \
+ dma_queue_push(dma_q, dma_make_data(dst->data, dst_spad), \
+ dst_row_size, dst_buf_size, dst_row_size, 0); \
+ dma_queue_push(dma_q, dma_make_data(src_spad, f_src0_addr), \
+ src0_buf_size, src0_row_size, src0_row_size, 1); \
+ f_src0_addr += nb01; \
+ if (++f_i01 == ne01) { \
+ f_i01 = 0; \
+ if (++f_i02 == ne02) { \
+ f_i02 = 0; \
+ f_i03++; \
+ } \
+ f_src0_addr = src0->data + f_i02*nb02 + f_i03*nb03; \
+ } \
+ } \
+ \
+ for (uint32_t r = 0; r < nrows_thread; ++r) { \
+ uint8_t * dst_spad = (uint8_t *) (uintptr_t) dma_queue_pop(dma_q).src; \
+ uint8_t * src_spad = (uint8_t *) (uintptr_t) dma_queue_pop(dma_q).dst; \
+ \
+ htp_trace_event_start(tr, HTP_TRACE_EVT_HVX_COMP, (uint16_t) r); \
+ CONV_FUNC(dst_spad, src_spad, ne00); \
+ htp_trace_event_stop(tr, HTP_TRACE_EVT_HVX_COMP, (uint16_t) r); \
+ \
+ dma_queue_push(dma_q, dma_make_data(c_dst_addr, dst_spad), \
+ dst_row_size, dst_buf_size, dst_row_size, 1); \
+ c_dst_addr += nb1; \
+ if (++c_i01 == ne01) { \
+ c_i01 = 0; \
+ if (++c_i02 == ne02) { \
+ c_i02 = 0; \
+ c_i03++; \
+ } \
+ c_dst_addr = dst->data + c_i02*nb2 + c_i03*nb3; \
+ } \
+ \
+ if (r + 2 < nrows_thread) { \
+ dma_queue_push(dma_q, dma_make_data(src_spad, f_src0_addr), \
+ src0_buf_size, src0_row_size, src0_row_size, 1); \
+ f_src0_addr += nb01; \
+ if (++f_i01 == ne01) { \
+ f_i01 = 0; \
+ if (++f_i02 == ne02) { \
+ f_i02 = 0; \
+ f_i03++; \
+ } \
+ f_src0_addr = src0->data + f_i02*nb02 + f_i03*nb03; \
+ } \
+ } \
+ } \
+ dma_queue_flush(dma_q); \
}
-static void cpy_thread_f32_f16_sameshape(unsigned int nth, unsigned int ith, void * data) {
- struct htp_copy_context * ct = (struct htp_copy_context *) data;
- struct htp_ops_context * octx = ct->octx;
- cpy_preamble;
-
- const uint32_t dr = ct->src0_nrows_per_thread;
- const uint32_t ir0 = ct->row_start + dr * ith;
- const uint32_t ir1 = MIN(ir0 + dr, ct->row_start + ct->nrows);
- if (ir0 >= ir1) return;
-
- const uint32_t ne02_ne01 = ne02 * ne01;
- uint32_t i03 = fastdiv(ir0, &ct->div_ne02_ne01);
- uint32_t rem = ir0 - i03 * ne02_ne01;
- uint32_t i02 = fastdiv(rem, &ct->div_ne01);
- uint32_t i01 = rem - i02 * ne01;
-
- uint8_t* dst_ptr = (uint8_t*) dst->data + i01*nb1 + i02*nb2 + i03*nb3;
- uint8_t* src0_ptr = (uint8_t*) src0->data + i01*nb01 + i02*nb02 + i03*nb03;
-
- for (uint32_t r = ir0; r < ir1; r++) {
- hex_l2fetch(src0_ptr, ne00 * sizeof(__fp16), nb01, 2);
- hvx_copy_f32_f16_uu(dst_ptr, src0_ptr, ne00);
- dst_ptr += nb1;
- src0_ptr += nb01;
- if (++i01 == ne01) {
- i01 = 0;
- if (++i02 == ne02) {
- i02 = 0;
- i03++;
- }
- dst_ptr = (uint8_t*) dst->data + i02*nb2 + i03*nb3;
- src0_ptr = (uint8_t*) src0->data + i02*nb02 + i03*nb03;
- }
+DEFINE_CPY_CONVERT_SAMESHAPE(f16_f32, hvx_copy_f16_f32_aa)
+DEFINE_CPY_CONVERT_SAMESHAPE(f32_f16, hvx_copy_f32_f16_aa)
+DEFINE_CPY_CONVERT_SAMESHAPE(i32_f32, hvx_copy_i32_f32_aa)
+DEFINE_CPY_CONVERT_SAMESHAPE(f32_i32, hvx_copy_f32_i32_aa)
+
+static int cpy_scalar(struct htp_ops_context * octx, const struct htp_copy_kernel_params * kparams) {
+ if (octx->ctx->mdev.count > 1 && octx->ctx->mdev.idx > 0) {
+ return HTP_STATUS_OK;
}
-}
-static void cpy_thread_i32_f32_sameshape(unsigned int nth, unsigned int ith, void * data) {
- struct htp_copy_context * ct = (struct htp_copy_context *) data;
- struct htp_ops_context * octx = ct->octx;
- cpy_preamble;
-
- const uint32_t dr = ct->src0_nrows_per_thread;
- const uint32_t ir0 = ct->row_start + dr * ith;
- const uint32_t ir1 = MIN(ir0 + dr, ct->row_start + ct->nrows);
- if (ir0 >= ir1) return;
-
- const uint32_t ne02_ne01 = ne02 * ne01;
- uint32_t i03 = fastdiv(ir0, &ct->div_ne02_ne01);
- uint32_t rem = ir0 - i03 * ne02_ne01;
- uint32_t i02 = fastdiv(rem, &ct->div_ne01);
- uint32_t i01 = rem - i02 * ne01;
-
- uint8_t* dst_ptr = (uint8_t*) dst->data + i01*nb1 + i02*nb2 + i03*nb3;
- uint8_t* src0_ptr = (uint8_t*) src0->data + i01*nb01 + i02*nb02 + i03*nb03;
-
- for (uint32_t r = ir0; r < ir1; r++) {
- hex_l2fetch(src0_ptr, ne00 * sizeof(float), nb01, 2);
- const float * restrict src_row = (const float *) src0_ptr;
- int32_t * restrict dst_row = (int32_t *) dst_ptr;
- for (uint32_t i = 0; i < ne00; i++) {
- dst_row[i] = (int32_t) src_row[i];
- }
- dst_ptr += nb1;
- src0_ptr += nb01;
- if (++i01 == ne01) {
- i01 = 0;
- if (++i02 == ne02) {
- i02 = 0;
- i03++;
- }
- dst_ptr = (uint8_t*) dst->data + i02*nb2 + i03*nb3;
- src0_ptr = (uint8_t*) src0->data + i02*nb02 + i03*nb03;
- }
+ const struct htp_tensor * src0 = octx->src[0];
+ const struct htp_tensor * dst = octx->dst;
+
+ if (src0->type == dst->type) {
+ dma_cpy_sametype_reshape_contig(octx->ctx->dma[0], dst->data, src0->data, kparams->src0_type_size);
+ dma_queue_flush(octx->ctx->dma[0]);
+ return HTP_STATUS_OK;
}
+
+ dma_queue * dma_q = octx->ctx->dma[0];
+ dma_addr_t s_vtcm = (dma_addr_t)(uintptr_t) octx->ctx->vtcm_base;
+ dma_addr_t d_vtcm = s_vtcm + VLEN;
+ const uint32_t s_size = kparams->src0_type_size;
+ const uint32_t d_size = kparams->dst_type_size;
+
+ dma_queue_push(dma_q, dma_make_data(s_vtcm, src0->data), s_size, s_size, s_size, 1);
+ dma_queue_pop(dma_q);
+
+ uint8_t * s_ptr = (uint8_t *) octx->ctx->vtcm_base;
+ uint8_t * d_ptr = s_ptr + VLEN;
+
+ const HVX_Vector v_src = hvx_vmem(s_ptr);
+ HVX_Vector v_dst;
+
+ if (src0->type == HTP_TYPE_F32 && dst->type == HTP_TYPE_I32) {
+ v_dst = Q6_Vw_equals_Vsf(v_src);
+ } else if (src0->type == HTP_TYPE_I32 && dst->type == HTP_TYPE_F32) {
+ v_dst = Q6_Vsf_equals_Vw(v_src);
+ } else if (src0->type == HTP_TYPE_F32 && dst->type == HTP_TYPE_F16) {
+ v_dst = hvx_vec_f32_to_f16(v_src, v_src);
+ } else if (src0->type == HTP_TYPE_F16 && dst->type == HTP_TYPE_F32) {
+ v_dst = Q6_V_lo_W(hvx_vec_f16_to_f32(v_src));
+ } else {
+ return HTP_STATUS_NO_SUPPORT;
+ }
+
+ hvx_vmem(d_ptr) = v_dst;
+
+ dma_queue_push(dma_q, dma_make_data(dst->data, d_vtcm), d_size, d_size, d_size, 1);
+ dma_queue_flush(dma_q);
+ return HTP_STATUS_OK;
}
-static void cpy_thread_f32_i32_sameshape(unsigned int nth, unsigned int ith, void * data) {
- struct htp_copy_context * ct = (struct htp_copy_context *) data;
- struct htp_ops_context * octx = ct->octx;
- cpy_preamble;
-
- const uint32_t dr = ct->src0_nrows_per_thread;
- const uint32_t ir0 = ct->row_start + dr * ith;
- const uint32_t ir1 = MIN(ir0 + dr, ct->row_start + ct->nrows);
- if (ir0 >= ir1) return;
-
- const uint32_t ne02_ne01 = ne02 * ne01;
- uint32_t i03 = fastdiv(ir0, &ct->div_ne02_ne01);
- uint32_t rem = ir0 - i03 * ne02_ne01;
- uint32_t i02 = fastdiv(rem, &ct->div_ne01);
- uint32_t i01 = rem - i02 * ne01;
-
- uint8_t* dst_ptr = (uint8_t*) dst->data + i01*nb1 + i02*nb2 + i03*nb3;
- uint8_t* src0_ptr = (uint8_t*) src0->data + i01*nb01 + i02*nb02 + i03*nb03;
-
- for (uint32_t r = ir0; r < ir1; r++) {
- hex_l2fetch(src0_ptr, ne00 * sizeof(int32_t), nb01, 2);
- const int32_t * restrict src_row = (const int32_t *) src0_ptr;
- float * restrict dst_row = (float *) dst_ptr;
- for (uint32_t i = 0; i < ne00; i++) {
- dst_row[i] = (float) src_row[i];
- }
- dst_ptr += nb1;
- src0_ptr += nb01;
- if (++i01 == ne01) {
- i01 = 0;
- if (++i02 == ne02) {
- i02 = 0;
- i03++;
- }
- dst_ptr = (uint8_t*) dst->data + i02*nb2 + i03*nb3;
- src0_ptr = (uint8_t*) src0->data + i02*nb02 + i03*nb03;
- }
+static int cpy_1d_contig(struct htp_ops_context * octx, const struct htp_copy_kernel_params * kparams) {
+ const struct htp_tensor * src0 = octx->src[0];
+ const struct htp_tensor * dst = octx->dst;
+
+ uint32_t elem_start = 0;
+ uint32_t nelem = kparams->total_elems;
+
+ if (octx->ctx->mdev.count > 1) {
+ const struct htp_tensor_mdev_range range = htp_tensor_mdev_partition(
+ kparams->total_elems, 1, octx->ctx->mdev.idx, octx->ctx->mdev.count, &octx->ctx->mdev.count_div);
+ elem_start = range.start;
+ nelem = range.count;
}
+
+ if (nelem > 0) {
+ dma_queue * q = octx->ctx->dma[0];
+ const uint32_t type_size = kparams->src0_type_size;
+ dma_addr_t dst_addr = dst->data + elem_start * type_size;
+ dma_addr_t src0_addr = src0->data + elem_start * type_size;
+ dma_cpy_sametype_reshape_contig(q, dst_addr, src0_addr, nelem * type_size);
+ dma_queue_flush(q);
+ }
+
+ return HTP_STATUS_OK;
}
-static int exec_cpy(struct htp_ops_context * octx, bool * use_dma) {
- cpy_preamble;
- *use_dma = false;
+static int cpy_sameshape_sametype(struct htp_ops_context * octx, const struct htp_copy_kernel_params * kparams) {
+ const struct htp_tensor * src0 = octx->src[0];
+ const struct htp_tensor * dst = octx->dst;
- const uint32_t total_elems_src = ne00 * ne01 * ne02 * ne03;
- const uint32_t total_elems_dst = ne0 * ne1 * ne2 * ne3;
- if (total_elems_src == 1 && total_elems_dst == 1) {
- if (octx->ctx->mdev.count > 1 && octx->ctx->mdev.idx > 0) {
- return HTP_STATUS_OK;
- }
- if (src0->type == HTP_TYPE_F32 && dst->type == HTP_TYPE_I32) {
- ((int32_t *) dst->data)[0] = (int32_t) (((const float *) src0->data)[0]);
- return HTP_STATUS_OK;
- }
- if (src0->type == HTP_TYPE_I32 && dst->type == HTP_TYPE_F32) {
- ((float *) dst->data)[0] = (float) (((const int32_t *) src0->data)[0]);
- return HTP_STATUS_OK;
- }
- if (src0->type == HTP_TYPE_I32 && dst->type == HTP_TYPE_I32) {
- ((int32_t *) dst->data)[0] = ((const int32_t *) src0->data)[0];
- return HTP_STATUS_OK;
- }
- if (src0->type == HTP_TYPE_F32 && dst->type == HTP_TYPE_F32) {
- ((float *) dst->data)[0] = ((const float *) src0->data)[0];
- return HTP_STATUS_OK;
- }
- if (src0->type == HTP_TYPE_F16 && dst->type == HTP_TYPE_F16) {
- ((__fp16 *) dst->data)[0] = ((const __fp16 *) src0->data)[0];
- return HTP_STATUS_OK;
- }
- if (src0->type == HTP_TYPE_F32 && dst->type == HTP_TYPE_F16) {
- ((__fp16 *) dst->data)[0] = (__fp16) (((const float *) src0->data)[0]);
- return HTP_STATUS_OK;
- }
- if (src0->type == HTP_TYPE_F16 && dst->type == HTP_TYPE_F32) {
- ((float *) dst->data)[0] = (float) (((const __fp16 *) src0->data)[0]);
- return HTP_STATUS_OK;
- }
+ uint32_t row_start = 0;
+ uint32_t nrows = kparams->total_rows;
+
+ if (octx->ctx->mdev.count > 1) {
+ const struct htp_tensor_mdev_range range = htp_tensor_mdev_partition(
+ kparams->total_rows, 1, octx->ctx->mdev.idx, octx->ctx->mdev.count, &octx->ctx->mdev.count_div);
+ row_start = range.start;
+ nrows = range.count;
}
- struct htp_copy_context ct;
- ct.octx = octx;
+ if (nrows > 0) {
+ dma_queue * q = octx->ctx->dma[0];
+ dma_cpy_sametype_sameshape_range(q, dst, src0, kparams->src0_type_size, row_start, nrows);
+ dma_queue_flush(q);
+ }
- switch (src0->type) {
- case HTP_TYPE_F32: ct.src0_type_size = 4; ct.src0_block_size = 1; ct.src0_blocks_per_row = ne00 / 1; break;
- case HTP_TYPE_F16: ct.src0_type_size = 2; ct.src0_block_size = 1; ct.src0_blocks_per_row = ne00 / 1; break;
- case HTP_TYPE_I32: ct.src0_type_size = 4; ct.src0_block_size = 1; ct.src0_blocks_per_row = ne00 / 1; break;
- default:
+ return HTP_STATUS_OK;
+}
+
+static int cpy_sameshape_convert(struct htp_ops_context * octx, const struct htp_copy_kernel_params * kparams) {
+ const struct htp_tensor * src0 = octx->src[0];
+ const struct htp_tensor * dst = octx->dst;
+
+ if (htp_tensor_is_extended(src0) || htp_tensor_is_extended(dst)) {
return HTP_STATUS_NO_SUPPORT;
}
- switch (dst->type) {
- case HTP_TYPE_F32: ct.dst_type_size = 4; ct.dst_block_size = 1; ct.dst_blocks_per_row = ne0 / 1; break;
- case HTP_TYPE_F16: ct.dst_type_size = 2; ct.dst_block_size = 1; ct.dst_blocks_per_row = ne0 / 1; break;
- case HTP_TYPE_I32: ct.dst_type_size = 4; ct.dst_block_size = 1; ct.dst_blocks_per_row = ne0 / 1; break;
- default:
- return HTP_STATUS_NO_SUPPORT;
+ if (!htp_ops_context_set_n_threads(octx, kparams->n_threads)) {
+ return HTP_STATUS_INVAL_PARAMS;
}
- const bool sametype = (src0->type == dst->type);
- const bool transposed = (nb00 > nb01) || (nb0 > nb1) ||
- (nb00 != ct.src0_type_size) || (nb0 != ct.dst_type_size) ||
- (nb01 < ne00 * ct.src0_type_size) || (nb1 < ne0 * ct.dst_type_size);
- const bool sameshape = !transposed && (ne00 == ne0 && ne01 == ne1 && ne02 == ne2 && ne03 == ne3);
+ uint32_t row_start = 0;
+ uint32_t nrows = kparams->total_rows;
- const uint32_t n_threads = octx->n_threads;
+ if (octx->ctx->mdev.count > 1) {
+ const struct htp_tensor_mdev_range range = htp_tensor_mdev_partition(
+ kparams->total_rows, 1, octx->ctx->mdev.idx, octx->ctx->mdev.count, &octx->ctx->mdev.count_div);
+ row_start = range.start;
+ nrows = range.count;
+ }
- const bool src_is_contiguous = htp_tensor_is_contiguous(src0, ct.src0_type_size);
- const bool dst_is_contiguous = htp_tensor_is_contiguous(dst, ct.dst_type_size);
+ if (nrows == 0) {
+ return HTP_STATUS_OK;
+ }
- if (htp_tensor_is_extended(src0) || htp_tensor_is_extended(dst)) {
- if (!sametype) {
- return HTP_STATUS_NO_SUPPORT;
- }
- if (!sameshape && !(src_is_contiguous && dst_is_contiguous && octx->ctx->mdev.count <= 1)) {
- return HTP_STATUS_NO_SUPPORT;
- }
+ if (kparams->vtcm_size > octx->ctx->vtcm_size) {
+ return HTP_STATUS_VTCM_TOO_SMALL;
}
- if (sameshape) {
- const uint32_t total_rows = ne01 * ne02 * ne03;
- const uint32_t row_size = ne00 * ct.dst_type_size;
+ const uint32_t n_threads = octx->n_threads;
+ const struct htp_copy_convert_params * cvt = &kparams->u.convert;
- ct.div_ne01 = init_fastdiv_values(ne01);
- ct.div_ne02_ne01 = init_fastdiv_values(ne02 * ne01);
+ struct htp_copy_context ct;
+ ct.octx = octx;
+ ct.kparams = kparams;
+ ct.row_start = row_start;
+ ct.nrows = nrows;
+ ct.src0_nrows_per_thread = fastdiv(nrows + n_threads - 1, &octx->n_threads_div);
+
+ uint8_t * vtcm_base = (uint8_t *) octx->ctx->vtcm_base;
+ ct.vtcm_src0 = vtcm_base;
+ ct.vtcm_dst = vtcm_base + (size_t) n_threads * cvt->spad0_size_per_thread;
+
+ work_queue_func_t copy_fun = NULL;
+ if (dst->type == HTP_TYPE_F16 && src0->type == HTP_TYPE_F32) {
+ copy_fun = cpy_thread_f16_f32_sameshape;
+ } else if (dst->type == HTP_TYPE_F32 && src0->type == HTP_TYPE_F16) {
+ copy_fun = cpy_thread_f32_f16_sameshape;
+ } else if (dst->type == HTP_TYPE_I32 && src0->type == HTP_TYPE_F32) {
+ copy_fun = cpy_thread_i32_f32_sameshape;
+ } else if (dst->type == HTP_TYPE_F32 && src0->type == HTP_TYPE_I32) {
+ copy_fun = cpy_thread_f32_i32_sameshape;
+ } else {
+ return HTP_STATUS_NO_SUPPORT;
+ }
- uint32_t row_start = 0;
- uint32_t nrows = total_rows;
+ work_queue_run(octx->ctx->work_queue, copy_fun, &ct, n_threads);
+ return HTP_STATUS_OK;
+}
- if (octx->ctx->mdev.count > 1) {
- const uint32_t rows_per_chunk = (row_size > 0) ? (HEX_L2_LINE_SIZE / hex_gcd_u32(row_size, HEX_L2_LINE_SIZE)) : 1;
- const bool can_split = htp_tensor_mdev_data_aligned(dst) && dst_is_contiguous;
- const struct htp_tensor_mdev_range range = htp_tensor_mdev_partition(total_rows, can_split ? rows_per_chunk : 0,
- octx->ctx->mdev.idx, octx->ctx->mdev.count, &octx->ctx->mdev.count_div);
- row_start = range.start;
- nrows = range.count;
- }
+static int cpy_reshape(struct htp_ops_context * octx, const struct htp_copy_kernel_params * kparams) {
+ const struct htp_tensor * src0 = octx->src[0];
+ const struct htp_tensor * dst = octx->dst;
- if (nrows == 0) {
- return HTP_STATUS_OK;
- }
+ if (htp_tensor_is_extended(src0) || htp_tensor_is_extended(dst)) {
+ return HTP_STATUS_NO_SUPPORT;
+ }
- ct.row_start = row_start;
- ct.nrows = nrows;
- ct.src0_nrows_per_thread = fastdiv(nrows + n_threads - 1, &octx->n_threads_div);
-
- if (sametype && (octx->ctx->mdev.count <= 1 || htp_tensor_is_extended(src0) || htp_tensor_is_extended(dst))) {
- if (octx->ctx->mdev.idx == 0) {
- *use_dma = true;
- cpy_dma_sametype_sameshape(octx->ctx->dma[0], dst, src0, ct.src0_type_size);
- dma_queue_flush(octx->ctx->dma[0]);
- }
- } else {
- work_queue_func_t copy_fun = NULL;
- if (sametype) {
- switch (src0->type) {
- case HTP_TYPE_F32: copy_fun = cpy_thread_f32_sameshape; break;
- case HTP_TYPE_F16: copy_fun = cpy_thread_f16_sameshape; break;
- case HTP_TYPE_I32: copy_fun = cpy_thread_i32_sameshape; break;
- default: return HTP_STATUS_NO_SUPPORT;
- }
- } else if (dst->type == HTP_TYPE_F16 && src0->type == HTP_TYPE_F32) {
- copy_fun = cpy_thread_f16_f32_sameshape;
- } else if (dst->type == HTP_TYPE_F32 && src0->type == HTP_TYPE_F16) {
- copy_fun = cpy_thread_f32_f16_sameshape;
- } else if (dst->type == HTP_TYPE_I32 && src0->type == HTP_TYPE_F32) {
- copy_fun = cpy_thread_i32_f32_sameshape;
- } else if (dst->type == HTP_TYPE_F32 && src0->type == HTP_TYPE_I32) {
- copy_fun = cpy_thread_f32_i32_sameshape;
- } else {
- return HTP_STATUS_NO_SUPPORT;
- }
- work_queue_run(octx->ctx->work_queue, copy_fun, &ct, n_threads);
- }
- } else if (sametype) {
- const uint32_t total_elems = ne0 * ne1 * ne2 * ne3;
- const uint32_t elems_per_line = (ct.dst_type_size == 4) ? 32 : 64;
-
- if (octx->ctx->mdev.count <= 1 && dst_is_contiguous && src_is_contiguous) {
- *use_dma = true;
- cpy_dma_sametype_reshape_contig(octx->ctx->dma[0], dst->data, src0->data, total_elems * ct.dst_type_size);
- dma_queue_flush(octx->ctx->dma[0]);
- return HTP_STATUS_OK;
- }
+ if (!htp_ops_context_set_n_threads(octx, kparams->n_threads)) {
+ return HTP_STATUS_INVAL_PARAMS;
+ }
- ct.div_ne0 = init_fastdiv_values(ne0);
- ct.div_ne1_ne0 = init_fastdiv_values(ne1 * ne0);
- ct.div_ne2_ne1_ne0 = init_fastdiv_values(ne2 * ne1 * ne0);
- ct.div_ne00 = init_fastdiv_values(ne00);
- ct.div_ne01_ne00 = init_fastdiv_values(ne01 * ne00);
- ct.div_ne02_ne01_ne00 = init_fastdiv_values(ne02 * ne01 * ne00);
-
- uint32_t elem_start = 0;
- uint32_t nelem = total_elems;
-
- if (octx->ctx->mdev.count > 1) {
- const bool can_split = htp_tensor_mdev_data_aligned(dst) && dst_is_contiguous;
- const struct htp_tensor_mdev_range range = htp_tensor_mdev_partition(total_elems, can_split ? elems_per_line : 0,
- octx->ctx->mdev.idx, octx->ctx->mdev.count, &octx->ctx->mdev.count_div);
- elem_start = range.start;
- nelem = range.count;
- }
+ uint32_t elem_start = 0;
+ uint32_t nelem = kparams->total_elems;
- if (nelem == 0) {
- return HTP_STATUS_OK;
- }
+ if (octx->ctx->mdev.count > 1) {
+ const struct htp_tensor_mdev_range range = htp_tensor_mdev_partition(
+ kparams->total_elems, 1, octx->ctx->mdev.idx, octx->ctx->mdev.count, &octx->ctx->mdev.count_div);
+ elem_start = range.start;
+ nelem = range.count;
+ }
- ct.elem_start = elem_start;
- ct.nelem = nelem;
- ct.elem_per_thread = fastdiv(nelem + n_threads - 1, &octx->n_threads_div);
+ if (nelem == 0) {
+ return HTP_STATUS_OK;
+ }
- work_queue_func_t copy_fun = NULL;
- switch (src0->type) {
- case HTP_TYPE_F32: copy_fun = cpy_thread_f32_reshape; break;
- case HTP_TYPE_F16: copy_fun = cpy_thread_f16_reshape; break;
- case HTP_TYPE_I32: copy_fun = cpy_thread_i32_reshape; break;
- default: return HTP_STATUS_NO_SUPPORT;
- }
- work_queue_run(octx->ctx->work_queue, copy_fun, &ct, n_threads);
- } else {
- return HTP_STATUS_NO_SUPPORT;
+ const uint32_t n_threads = octx->n_threads;
+
+ struct htp_copy_context ct;
+ ct.octx = octx;
+ ct.kparams = kparams;
+ ct.elem_start = elem_start;
+ ct.nelem = nelem;
+ ct.elem_per_thread = fastdiv(nelem + n_threads - 1, &octx->n_threads_div);
+
+ work_queue_func_t copy_fun = NULL;
+ switch (src0->type) {
+ case HTP_TYPE_F32: copy_fun = cpy_thread_f32_reshape; break;
+ case HTP_TYPE_F16: copy_fun = cpy_thread_f16_reshape; break;
+ case HTP_TYPE_I32: copy_fun = cpy_thread_i32_reshape; break;
+ default: return HTP_STATUS_NO_SUPPORT;
}
+ work_queue_run(octx->ctx->work_queue, copy_fun, &ct, n_threads);
return HTP_STATUS_OK;
}
int op_cpy(struct htp_ops_context * octx) {
- bool use_dma = false;
- int status = exec_cpy(octx, &use_dma);
+ const struct htp_copy_kernel_params * kparams = (const struct htp_copy_kernel_params *) octx->kernel_params;
+ int status = HTP_STATUS_OK;
+
+ switch (kparams->kernel_type) {
+ case HTP_COPY_KERNEL_SCALAR:
+ status = cpy_scalar(octx, kparams);
+ break;
+ case HTP_COPY_KERNEL_1D_CONTIG:
+ status = cpy_1d_contig(octx, kparams);
+ break;
+ case HTP_COPY_KERNEL_SAMESHAPE_SAMETYPE:
+ status = cpy_sameshape_sametype(octx, kparams);
+ break;
+ case HTP_COPY_KERNEL_SAMESHAPE_CONVERT:
+ status = cpy_sameshape_convert(octx, kparams);
+ break;
+ case HTP_COPY_KERNEL_RESHAPE:
+ status = cpy_reshape(octx, kparams);
+ break;
+ default:
+ status = HTP_STATUS_NO_SUPPORT;
+ break;
+ }
htp_ops_context_set_status(octx, status);
- if (octx->op == HTP_OP_CPY_FENCE) {
- if (!use_dma) {
- htp_flush_dirty_ranges(octx->ctx);
- }
-
+ if (octx->ctx->mdev.count > 1) {
htp_mdev_group_barrier(octx);
+ }
+ if (octx->op == HTP_OP_CPY_FENCE) {
if (octx->ctx->mdev.idx == 0) {
const struct htp_tensor * sync = octx->src[1];
- if (htp_tensor_is_extended(sync)) {
- return HTP_STATUS_NO_SUPPORT;
- }
const uint32_t seq = (uint32_t) octx->op_params[0];
atomic_uint * sync_fence = (atomic_uint *) (uintptr_t) sync->data;
htp_fence_write(sync_fence, seq, octx->status);
diff --git a/ggml/src/ggml-hexagon/htp/cpy-ops.h b/ggml/src/ggml-hexagon/htp/cpy-ops.h
new file mode 100644
index 000000000..4cf7dc13b
--- /dev/null
+++ b/ggml/src/ggml-hexagon/htp/cpy-ops.h
@@ -0,0 +1,79 @@
+#ifndef HTP_CPY_OPS_H
+#define HTP_CPY_OPS_H
+
+#include "hex-common.h"
+#include "hex-fastdiv.h"
+#include <stdint.h>
+
+enum htp_copy_kernel_type {
+ HTP_COPY_KERNEL_UNSUPPORTED = 0,
+ HTP_COPY_KERNEL_1D_CONTIG = 1,
+ HTP_COPY_KERNEL_SAMESHAPE_SAMETYPE = 2,
+ HTP_COPY_KERNEL_SAMESHAPE_CONVERT = 3,
+ HTP_COPY_KERNEL_RESHAPE = 4,
+ HTP_COPY_KERNEL_SCALAR = 5,
+};
+
+struct htp_copy_convert_params {
+ uint32_t src0_buf_size;
+ uint32_t dst_buf_size;
+ uint32_t spad0_size_per_thread;
+ uint32_t spad1_size_per_thread;
+ struct fastdiv_values div_ne01;
+ struct fastdiv_values div_ne02_ne01;
+};
+
+struct htp_copy_reshape_params {
+ struct fastdiv_values div_ne0;
+ struct fastdiv_values div_ne1_ne0;
+ struct fastdiv_values div_ne2_ne1_ne0;
+ struct fastdiv_values div_ne00;
+ struct fastdiv_values div_ne01_ne00;
+ struct fastdiv_values div_ne02_ne01_ne00;
+};
+
+struct htp_copy_kernel_params {
+ uint8_t kernel_type;
+ uint8_t src0_type_size;
+ uint8_t dst_type_size;
+ uint8_t n_threads;
+
+ uint32_t total_elems;
+ uint32_t total_rows;
+ uint32_t vtcm_size;
+
+ union {
+ struct htp_copy_convert_params convert;
+ struct htp_copy_reshape_params reshape;
+ } u;
+};
+
+struct htp_copy_convert_vtcm_layout {
+ uint32_t src0_buf_size;
+ uint32_t dst_buf_size;
+ uint32_t spad0_size_per_thread;
+ uint32_t spad1_size_per_thread;
+ uint32_t total_bytes;
+};
+
+static inline void htp_copy_convert_vtcm_layout_build(
+ struct htp_copy_convert_vtcm_layout * layout,
+ uint32_t ne00,
+ uint32_t src_type_size,
+ uint32_t dst_type_size,
+ uint32_t n_threads) {
+
+ layout->src0_buf_size = hex_round_up(ne00 * src_type_size, 256);
+ layout->dst_buf_size = hex_round_up(ne00 * dst_type_size, 256);
+ layout->spad0_size_per_thread = 2 * layout->src0_buf_size;
+ layout->spad1_size_per_thread = 2 * layout->dst_buf_size;
+ layout->total_bytes = n_threads * (layout->spad0_size_per_thread + layout->spad1_size_per_thread);
+}
+
+#if defined(__cplusplus)
+static_assert(sizeof(struct htp_copy_kernel_params) <= 128, "htp_copy_kernel_params is too large for kernel_params blob");
+#else
+_Static_assert(sizeof(struct htp_copy_kernel_params) <= 128, "htp_copy_kernel_params is too large for kernel_params blob");
+#endif
+
+#endif // HTP_CPY_OPS_H
diff --git a/ggml/src/ggml-hexagon/htp/hex-cpy-dma.h b/ggml/src/ggml-hexagon/htp/dma-copy.h
similarity index 58%
rename from ggml/src/ggml-hexagon/htp/hex-cpy-dma.h
rename to ggml/src/ggml-hexagon/htp/dma-copy.h
index c87f495e2..b2c91a13c 100644
--- a/ggml/src/ggml-hexagon/htp/hex-cpy-dma.h
+++ b/ggml/src/ggml-hexagon/htp/dma-copy.h
@@ -1,9 +1,9 @@
-#ifndef HEX_CPY_DMA_H
-#define HEX_CPY_DMA_H
+#ifndef HTP_DMA_COPY_H
+#define HTP_DMA_COPY_H
// DDR<->DDR DMA copies of same-type, same-shape tensors with arbitrary strides.
// Used by CPY for the copy itself and by CONCAT, which is two such copies into
-// two views of its destination. Every helper only pushes descriptors; the
+// two views of its destination. Every helper only pushes descriptors; the
// caller flushes the queue when it needs the data.
#include "dma-queue.h"
@@ -14,7 +14,7 @@
#include <stdint.h>
// Contiguous byte run, as 1d transfers of at most DMA_SAFE_CHUNK_SIZE each.
-static inline void cpy_dma_sametype_reshape_contig(dma_queue * dma_q,
+static inline void dma_cpy_sametype_reshape_contig(dma_queue * dma_q,
dma_addr_t dst,
dma_addr_t src0,
uint32_t total_bytes) {
@@ -36,7 +36,7 @@ static inline void cpy_dma_sametype_reshape_contig(dma_queue * dma_q,
}
// One 2d transfer, split at the 16-bit nrows field.
-static inline void cpy_dma_push_2d_chunked(dma_queue * dma_q,
+static inline void dma_cpy_push_2d_chunked(dma_queue * dma_q,
dma_addr_t dst,
dma_addr_t src,
size_t dst_stride,
@@ -59,12 +59,18 @@ static inline void cpy_dma_push_2d_chunked(dma_queue * dma_q,
}
}
-// Copy src0 into dst: same type, same ne[], any nb[] above dim 0, dim 0 dense on
-// both sides (nb[0] == elem_size).
-static inline void cpy_dma_sametype_sameshape(dma_queue * dma_q,
- const struct htp_tensor * dst,
- const struct htp_tensor * src0,
- uint32_t elem_size) {
+// Copy a range of rows [row_start, row_start + nrows) from src0 into dst:
+// same type, same ne[], any nb[] above dim 0, dim 0 dense on both sides (nb[0] == elem_size).
+static inline void dma_cpy_sametype_sameshape_range(dma_queue * dma_q,
+ const struct htp_tensor * dst,
+ const struct htp_tensor * src0,
+ uint32_t elem_size,
+ uint32_t row_start,
+ uint32_t nrows) {
+ if (nrows == 0) {
+ return;
+ }
+
const uint32_t ne00 = src0->ne[0];
const uint32_t ne01 = src0->ne[1];
const uint32_t ne02 = src0->ne[2];
@@ -85,13 +91,16 @@ static inline void cpy_dma_sametype_sameshape(dma_queue * dma_q,
const bool contiguous = htp_tensor_is_contiguous(src0, elem_size) && htp_tensor_is_contiguous(dst, elem_size);
if (contiguous) {
- cpy_dma_sametype_reshape_contig(dma_q, dst->data, src0->data, ne00 * elem_size * ne01 * ne02 * ne03);
+ dma_cpy_sametype_reshape_contig(dma_q,
+ dst->data + (dma_addr_t) row_start * ne00 * elem_size,
+ src0->data + (dma_addr_t) row_start * ne00 * elem_size,
+ nrows * ne00 * elem_size);
return;
}
// The single-descriptor path flattens (i01,i02,i03) into one row index, so every
// row must sit at a constant stride: nb01 on the source, nb1 on the destination.
- // Walk the outer dims and require each to continue that progression. A dim of
+ // Walk the outer dims and require each to continue that progression. A dim of
// extent 1 spans no rows, so it is skipped -- but its own stride must NOT then be
// used to justify the next dim's stride, which is what comparing nb03 against
// ne02*nb02 did: ggml leaves the stride of an extent-1 dim meaningless, so a view
@@ -109,18 +118,52 @@ static inline void cpy_dma_sametype_sameshape(dma_queue * dma_q,
}
if (contiguous_outer) {
- uint32_t total_rows = ne01 * ne02 * ne03;
- cpy_dma_push_2d_chunked(dma_q, dst->data, src0->data, nb1, nb01, ne00 * elem_size, total_rows);
+ dma_cpy_push_2d_chunked(dma_q,
+ dst->data + (dma_addr_t) row_start * nb1,
+ src0->data + (dma_addr_t) row_start * nb01,
+ nb1, nb01, ne00 * elem_size, nrows);
return;
}
- for (uint32_t i03 = 0; i03 < ne03; i03++) {
- for (uint32_t i02 = 0; i02 < ne02; i02++) {
- dma_addr_t dst_data = dst->data + i02 * nb2 + i03 * nb3;
- dma_addr_t src0_data = src0->data + i02 * nb02 + i03 * nb03;
- cpy_dma_push_2d_chunked(dma_q, dst_data, src0_data, nb1, nb01, ne00 * elem_size, ne01);
+ const uint32_t ne02_ne01 = ne02 * ne01;
+ uint32_t i03 = row_start / ne02_ne01;
+ uint32_t rem = row_start - i03 * ne02_ne01;
+ uint32_t i02 = rem / ne01;
+ uint32_t i01 = rem - i02 * ne01;
+
+ dma_addr_t cur_dst = dst->data + (dma_addr_t) i01 * nb1 + (dma_addr_t) i02 * nb2 + (dma_addr_t) i03 * nb3;
+ dma_addr_t cur_src0 = src0->data + (dma_addr_t) i01 * nb01 + (dma_addr_t) i02 * nb02 + (dma_addr_t) i03 * nb03;
+
+ uint32_t r = row_start;
+ const uint32_t row_end = row_start + nrows;
+ while (r < row_end) {
+ uint32_t cur_rows = MIN(row_end - r, ne01 - i01);
+ dma_cpy_push_2d_chunked(dma_q, cur_dst, cur_src0, nb1, nb01, ne00 * elem_size, cur_rows);
+ r += cur_rows;
+ i01 += cur_rows;
+ if (i01 == ne01) {
+ i01 = 0;
+ if (++i02 == ne02) {
+ i02 = 0;
+ i03++;
+ }
+ cur_dst = dst->data + (dma_addr_t) i02 * nb2 + (dma_addr_t) i03 * nb3;
+ cur_src0 = src0->data + (dma_addr_t) i02 * nb02 + (dma_addr_t) i03 * nb03;
+ } else {
+ cur_dst += cur_rows * nb1;
+ cur_src0 += cur_rows * nb01;
}
}
}
-#endif /* HEX_CPY_DMA_H */
+// Copy src0 into dst: same type, same ne[], any nb[] above dim 0, dim 0 dense on
+// both sides (nb[0] == elem_size).
+static inline void dma_cpy_sametype_sameshape(dma_queue * dma_q,
+ const struct htp_tensor * dst,
+ const struct htp_tensor * src0,
+ uint32_t elem_size) {
+ const uint32_t total_rows = src0->ne[1] * src0->ne[2] * src0->ne[3];
+ dma_cpy_sametype_sameshape_range(dma_q, dst, src0, elem_size, 0, total_rows);
+}
+
+#endif /* HTP_DMA_COPY_H */
diff --git a/ggml/src/ggml-hexagon/htp/hex-utils.h b/ggml/src/ggml-hexagon/htp/hex-utils.h
index 853f1c1b2..abb44da25 100644
--- a/ggml/src/ggml-hexagon/htp/hex-utils.h
+++ b/ggml/src/ggml-hexagon/htp/hex-utils.h
@@ -38,21 +38,15 @@ static inline void hex_l2fetch_block(const void * addr, size_t size) {
}
#define HEX_L2_LINE_SIZE 128
-#define HEX_L2_BLOCK_SIZE (HEX_L2_LINE_SIZE * 4) // flush granularity (lines per loop iteration)
+#define HEX_L2_BLOCK_SIZE (HEX_L2_LINE_SIZE * 4) // flush granularity (chunks per thread)
#define HEX_L2_FLUSH_WQ_THRESHOLD (4 * 1024)
#define HEX_L2_FLUSH_ALL_THRESHOLD (4 * 1024 * 1024)
static inline void hex_l2flush(void * addr, size_t size) {
+ if (size == 0) return;
const uint32_t s = ((uint32_t) addr) & ~(HEX_L2_LINE_SIZE - 1);
const uint32_t e = (((uint32_t) addr) + size + HEX_L2_LINE_SIZE - 1) & ~(HEX_L2_LINE_SIZE - 1);
- const uint32_t eb = s + ((e - s) & ~(HEX_L2_BLOCK_SIZE - 1));
- for (uint32_t i = s; i < eb; i += HEX_L2_BLOCK_SIZE) {
- Q6_dccleaninva_A((void *) (i + HEX_L2_LINE_SIZE * 0));
- Q6_dccleaninva_A((void *) (i + HEX_L2_LINE_SIZE * 1));
- Q6_dccleaninva_A((void *) (i + HEX_L2_LINE_SIZE * 2));
- Q6_dccleaninva_A((void *) (i + HEX_L2_LINE_SIZE * 3));
- }
- for (uint32_t i = eb; i < e; i += HEX_L2_LINE_SIZE) {
+ for (uint32_t i = s; i < e; i += HEX_L2_LINE_SIZE) {
Q6_dccleaninva_A((void *) i);
}
}
diff --git a/ggml/src/ggml-hexagon/htp/hvx-copy.h b/ggml/src/ggml-hexagon/htp/hvx-copy.h
index a3e33c3b3..08eb7cd61 100644
--- a/ggml/src/ggml-hexagon/htp/hvx-copy.h
+++ b/ggml/src/ggml-hexagon/htp/hvx-copy.h
@@ -44,7 +44,7 @@ static inline void hvx_splat_f32_u(void * restrict dst, float v, uint32_t n) {
}
static inline void hvx_splat_f16_a(void * restrict dst, _Float16 v, uint32_t n) {
- hvx_splat_u(dst, hvx_vec_splat_f16(v), n, sizeof(__fp16));
+ hvx_splat_a(dst, hvx_vec_splat_f16(v), n, sizeof(__fp16));
}
static inline void hvx_splat_f16_u(void * restrict dst, _Float16 v, uint32_t n) {
@@ -106,42 +106,42 @@ static inline void hvx_copy_uu(uint8_t * restrict dst, const uint8_t * restrict
hvx_copy_loop_body(HVX_UVector, HVX_UVector, hvx_vec_store_u);
}
-// copy n fp16 elements : source and destination are aligned to HVX Vector (128)
+// copy n fp16 elements : destination and source are aligned to HVX Vector (128)
static inline void hvx_copy_f16_aa(uint8_t * restrict dst, const uint8_t * restrict src, uint32_t n) {
hvx_copy_aa(dst, src, n, sizeof(__fp16));
}
-// copy n fp16 elements : source is aligned, destination is potentially unaligned
+// copy n fp16 elements : destination is aligned, source is unaligned
static inline void hvx_copy_f16_au(uint8_t * restrict dst, const uint8_t * restrict src, uint32_t n) {
hvx_copy_au(dst, src, n, sizeof(__fp16));
}
-// copy n fp16 elements : source is aligned, destination is potentially unaligned
+// copy n fp16 elements : destination is unaligned, source is aligned
static inline void hvx_copy_f16_ua(uint8_t * restrict dst, const uint8_t * restrict src, uint32_t n) {
hvx_copy_ua(dst, src, n, sizeof(__fp16));
}
-// copy n fp16 elements : source is aligned, destination is potentially unaligned
+// copy n fp16 elements : destination and source are unaligned
static inline void hvx_copy_f16_uu(uint8_t * restrict dst, const uint8_t * restrict src, uint32_t n) {
hvx_copy_uu(dst, src, n, sizeof(__fp16));
}
-// copy n fp32 elements : source and destination are aligned to HVX Vector (128)
+// copy n fp32 elements : destination and source are aligned to HVX Vector (128)
static inline void hvx_copy_f32_aa(uint8_t * restrict dst, const uint8_t * restrict src, uint32_t n) {
hvx_copy_aa(dst, src, n, sizeof(float));
}
-// copy n fp32 elements : source is aligned, destination is unaligned
-static inline void hvx_copy_f32_ua(uint8_t * restrict dst, const uint8_t * restrict src, uint32_t n) {
- hvx_copy_ua(dst, src, n, sizeof(float));
-}
-
-// copy n fp32 elements : source is unaligned, destination is aligned
+// copy n fp32 elements : destination is aligned, source is unaligned
static inline void hvx_copy_f32_au(uint8_t * restrict dst, const uint8_t * restrict src, uint32_t n) {
hvx_copy_au(dst, src, n, sizeof(float));
}
-// copy n fp32 elements : source is unaligned, destination unaligned
+// copy n fp32 elements : destination is unaligned, source is aligned
+static inline void hvx_copy_f32_ua(uint8_t * restrict dst, const uint8_t * restrict src, uint32_t n) {
+ hvx_copy_ua(dst, src, n, sizeof(float));
+}
+
+// copy n fp32 elements : destination and source are unaligned
static inline void hvx_copy_f32_uu(uint8_t * restrict dst, const uint8_t * restrict src, uint32_t n) {
hvx_copy_uu(dst, src, n, sizeof(float));
}
@@ -170,26 +170,26 @@ static inline void hvx_copy_f32_uu(uint8_t * restrict dst, const uint8_t * restr
} \
} while(0)
-// copy/convert n fp32 elements into n fp16 elements : source is aligned, destination is aligned
+// copy/convert n fp32 elements into n fp16 elements : destination and source are aligned
static inline void hvx_copy_f16_f32_aa(uint8_t * restrict dst, const uint8_t * restrict src, uint32_t n) {
assert((unsigned long) dst % 128 == 0);
assert((unsigned long) src % 128 == 0);
hvx_copy_f16_f32_loop_body(HVX_Vector, HVX_Vector, hvx_vec_store_a);
}
-// copy/convert n fp32 elements into n fp16 elements : source is unaligned, destination is aligned
+// copy/convert n fp32 elements into n fp16 elements : destination is aligned, source is unaligned
static inline void hvx_copy_f16_f32_au(uint8_t * restrict dst, const uint8_t * restrict src, uint32_t n) {
assert((unsigned long) dst % 128 == 0);
hvx_copy_f16_f32_loop_body(HVX_Vector, HVX_UVector, hvx_vec_store_a);
}
-// copy/convert n fp32 elements into n fp16 elements : source is aligned, destination is unaligned
+// copy/convert n fp32 elements into n fp16 elements : destination is unaligned, source is aligned
static inline void hvx_copy_f16_f32_ua(uint8_t * restrict dst, const uint8_t * restrict src, uint32_t n) {
assert((unsigned long) src % 128 == 0);
hvx_copy_f16_f32_loop_body(HVX_UVector, HVX_Vector, hvx_vec_store_u);
}
-// copy/convert n fp32 elements into n fp16 elements : source is unaligned, destination is unaligned
+// copy/convert n fp32 elements into n fp16 elements : destination and source are unaligned
static inline void hvx_copy_f16_f32_uu(uint8_t * restrict dst, const uint8_t * restrict src, uint32_t n) {
hvx_copy_f16_f32_loop_body(HVX_UVector, HVX_UVector, hvx_vec_store_u);
}
@@ -235,28 +235,98 @@ static inline void hvx_copy_f16_f32_uu(uint8_t * restrict dst, const uint8_t * r
} \
} while(0)
-// copy/convert n fp16 elements into n fp32 elements : source is aligned, destination is aligned
+// copy/convert n fp16 elements into n fp32 elements : destination and source are aligned
static inline void hvx_copy_f32_f16_aa(uint8_t * restrict dst, const uint8_t * restrict src, uint32_t n) {
assert((unsigned long) dst % 128 == 0);
assert((unsigned long) src % 128 == 0);
hvx_copy_f32_f16_loop_body(HVX_Vector, HVX_Vector, hvx_vec_store_a);
}
-// copy/convert n fp16 elements into n fp32 elements : source is unaligned, destination is aligned
+// copy/convert n fp16 elements into n fp32 elements : destination is aligned, source is unaligned
static inline void hvx_copy_f32_f16_au(uint8_t * restrict dst, const uint8_t * restrict src, uint32_t n) {
assert((unsigned long) dst % 128 == 0);
hvx_copy_f32_f16_loop_body(HVX_Vector, HVX_UVector, hvx_vec_store_a);
}
-// copy/convert n fp16 elements into n fp32 elements : source is aligned, destination is unaligned
+// copy/convert n fp16 elements into n fp32 elements : destination is unaligned, source is aligned
static inline void hvx_copy_f32_f16_ua(uint8_t * restrict dst, const uint8_t * restrict src, uint32_t n) {
assert((unsigned long) src % 128 == 0);
hvx_copy_f32_f16_loop_body(HVX_UVector, HVX_Vector, hvx_vec_store_u);
}
-// copy/convert n fp16 elements into n fp32 elements : source is unaligned, destination is unaligned
+// copy/convert n fp16 elements into n fp32 elements : destination and source are unaligned
static inline void hvx_copy_f32_f16_uu(uint8_t * restrict dst, const uint8_t * restrict src, uint32_t n) {
hvx_copy_f32_f16_loop_body(HVX_UVector, HVX_UVector, hvx_vec_store_u);
}
+//// fp32 -> int32
+
+#define hvx_copy_i32_f32_loop_body(dst_type, src_type, vec_store) \
+ do { \
+ dst_type * restrict vdst = (dst_type *) dst; \
+ src_type * restrict vsrc = (src_type *) src; \
+ \
+ const uint32_t elem_size = sizeof(int32_t); \
+ const uint32_t epv = 128 / elem_size; \
+ const uint32_t nvec = n / epv; \
+ const uint32_t nloe = n % epv; \
+ \
+ uint32_t i = 0; \
+ _Pragma("unroll(4)") \
+ for (; i < nvec; i++) { \
+ vdst[i] = Q6_Vw_equals_Vsf(vsrc[i]); \
+ } \
+ if (nloe) { \
+ HVX_Vector v = Q6_Vw_equals_Vsf(vsrc[i]); \
+ vec_store((void *) &vdst[i], nloe * elem_size, v); \
+ } \
+ } while(0)
+
+// copy/convert n fp32 elements into n int32 elements : destination and source are aligned
+static inline void hvx_copy_i32_f32_aa(uint8_t * restrict dst, const uint8_t * restrict src, uint32_t n) {
+ assert((unsigned long) dst % 128 == 0);
+ assert((unsigned long) src % 128 == 0);
+ hvx_copy_i32_f32_loop_body(HVX_Vector, HVX_Vector, hvx_vec_store_a);
+}
+
+// copy/convert n fp32 elements into n int32 elements : destination and source are unaligned
+static inline void hvx_copy_i32_f32_uu(uint8_t * restrict dst, const uint8_t * restrict src, uint32_t n) {
+ hvx_copy_i32_f32_loop_body(HVX_UVector, HVX_UVector, hvx_vec_store_u);
+}
+
+//// int32 -> fp32
+
+#define hvx_copy_f32_i32_loop_body(dst_type, src_type, vec_store) \
+ do { \
+ dst_type * restrict vdst = (dst_type *) dst; \
+ src_type * restrict vsrc = (src_type *) src; \
+ \
+ const uint32_t elem_size = sizeof(float); \
+ const uint32_t epv = 128 / elem_size; \
+ const uint32_t nvec = n / epv; \
+ const uint32_t nloe = n % epv; \
+ \
+ uint32_t i = 0; \
+ _Pragma("unroll(4)") \
+ for (; i < nvec; i++) { \
+ vdst[i] = Q6_Vsf_equals_Vw(vsrc[i]); \
+ } \
+ if (nloe) { \
+ HVX_Vector v = Q6_Vsf_equals_Vw(vsrc[i]); \
+ vec_store((void *) &vdst[i], nloe * elem_size, v); \
+ } \
+ } while(0)
+
+// copy/convert n int32 elements into n fp32 elements : destination and source are aligned
+static inline void hvx_copy_f32_i32_aa(uint8_t * restrict dst, const uint8_t * restrict src, uint32_t n) {
+ assert((unsigned long) dst % 128 == 0);
+ assert((unsigned long) src % 128 == 0);
+ hvx_copy_f32_i32_loop_body(HVX_Vector, HVX_Vector, hvx_vec_store_a);
+}
+
+// copy/convert n int32 elements into n fp32 elements : destination and source are unaligned
+static inline void hvx_copy_f32_i32_uu(uint8_t * restrict dst, const uint8_t * restrict src, uint32_t n) {
+ hvx_copy_f32_i32_loop_body(HVX_UVector, HVX_UVector, hvx_vec_store_u);
+}
+
#endif // HVX_COPY_H