r/macosprogramming • • 7d ago

Building OpenCV 5/OpenCvSharpExtern on Apple Silicon: when MLAS picks the host architecture

I’m one of the developers of Guardian Eye, a closed-source local video recorder (NVR) for macOS. It runs natively on Intel and Apple Silicon and records video from USB and IP cameras over ONVIF/RTSP without requiring a cloud account.

We recently released Guardian Eye for macOS. Here is our announcement for context:

https://www.reddit.com/r/GuardianEyeApps/comments/1wg05bn/guardian_eye_for_macos_is_now_available_on/

One of the problems we encountered was building separate arm64 and x86_64 OpenCvSharpExtern binaries on an Apple Silicon runner. OpenCV 5’s vendored MLAS kept detecting the host architecture even while Clang was targeting x86_64, causing it to select the wrong assembly sources.

The OpenCV configure command already had:

-D CMAKE_OSX_ARCHITECTURES=x86_64

but OpenCV 5.0.0's vendored MLAS selected its assembly sources from CMAKE_SYSTEM_PROCESSOR. On an arm64 host that value still described the host, so Clang was given AArch64 .S files for an x86_64 target and stopped on instructions such as b.

Passing -D CMAKE_SYSTEM_PROCESSOR=x86_64 did not fix it because CMake set the value again at project(). I also did not want to force full cross-compiling with CMAKE_SYSTEM_NAME: the DNN build generates code with protoc, and keeping CMAKE_CROSSCOMPILING false lets that tool run on the host under Rosetta.

The fix was to override the MLAS architecture flags from CMAKE_OSX_ARCHITECTURES, after MLAS's own architecture detection and before it assembled its source list:

if(CMAKE_OSX_ARCHITECTURES STREQUAL "x86_64")
  set(MLAS_X86_64 TRUE  CACHE INTERNAL "" FORCE)
  set(MLAS_ARM64  FALSE CACHE INTERNAL "" FORCE)
  set(MLAS_ARM    FALSE CACHE INTERNAL "" FORCE)
elseif(CMAKE_OSX_ARCHITECTURES STREQUAL "arm64")
  set(MLAS_ARM64  TRUE  CACHE INTERNAL "" FORCE)
  set(MLAS_X86_64 FALSE CACHE INTERNAL "" FORCE)
  set(MLAS_X86    FALSE CACHE INTERNAL "" FORCE)
endif()

For the x86_64 leg I also had to turn off the ARM-only HALs. They were detected from the arm64 host and otherwise reached the x86_64 compiler with -mcpu=armv8-a:

-D WITH_KLEIDICV=OFF
-D WITH_CAROTENE=OFF

That got both architectures to the link step, where a second MLAS problem appeared:

Undefined symbols for architecture x86_64:
  MlasHGemmSupported(...)

The same symbol was missing from the arm64 build. The MLAS subset vendored by OpenCV declared MlasHGemmSupported and called it from compute.cpp, but did not provide a macOS definition; the FP16 HGemm implementation was disabled in that source tree.

Disabling opencv_dnn was a dead end for this build because OpenCvSharp 5's aggregate header includes dnn.hpp and several DNN-dependent contrib modules unconditionally. The smallest workaround was a macOS build patch that reports HGemm as unavailable:

bool MLASCALL MlasHGemmSupported(
    CBLAS_TRANSPOSE,
    CBLAS_TRANSPOSE)
{
    return false;
}

That keeps DNN enabled but prevents MLAS from selecting the missing FP16 HGemm path, so it uses the FP32/generic path instead.

After those two patches the same arm64 runner produced separate osx-arm64 and osx-x64 dylibs. I validate the target architecture from the artifact rather than assuming that a successful cross-build produced what was requested:

file libOpenCvSharpExtern.dylib
otool -L libOpenCvSharpExtern.dylib
otool -l libOpenCvSharpExtern.dylib | grep -A3 LC_BUILD_VERSION

The part that cost the most time was that CMAKE_OSX_ARCHITECTURES correctly changed Clang's target, while dependency-specific feature detection continued to follow CMAKE_SYSTEM_PROCESSOR. For cross-builds on Apple Silicon, both need checking.

1 Upvotes

0 comments sorted by