Resumen de: US20260241941A1
A computer-implemented method for monitoring an artificial deep neural network comprises supplying input data to the trained deep neural network monitored, in order to obtain therefrom output data activation map data, and supplying the input data, the output data and the activation map data to a computer-implemented network observer. The network observer generates masking data from the activation map data and/or the input data and/or the output data; masks the activation map data using the masking data in order to obtain masked activation map data, wherein the masked activation map data contain unmasked activation values and masked activation values; and determines an outlier score for the output data using the masked activation map data, wherein merely the unmasked values are taken into account when determining the outlier score. The outlier score is a numerical value and indicates the extent to which the determined output deviates from a typical case.
Resumen de: WO2026170664A1
The present invention relates to the technical field of long-tailed medical image classification. Disclosed are a contrastive-learning-based medical image classification method and system, and a storage medium. The method comprises: dividing a long-tailed medical image data set into a training set and a test set, and on the basis of a preset scheme and in the form of batch processing, separately performing weak data augmentation and strong data augmentation on images in the training set; by means of a deep convolutional neural network, performing a contrastive learning task on obtained weakly data-augmented images and obtained strongly data-augmented images, and learning network parameters to obtain a parameter-optimized deep convolutional neural network; and classifying long-tailed medical images in the test set. In the present application, by means of a prototype-enhanced contrastive learning strategy, learnable-class prototypes are generated and subjected to data augmentation, so as to obtain a balanced implicitly-augmented contrastive learning loss, thereby achieving the advantages of high precision, high efficiency, low costs, wide applicability, etc.
Resumen de: US20260246928A1
0000 Systems and methods are provided for encoding and decoding video for machine consumption in which bandwidth is reduced by filtering feature layers and filtering channels at the encoder site that are determined to be redundant or of reduced relevance. A video encoder includes a neural network front end which receives image data and generates a plurality of feature layers. The relevance of the plurality of feature layers to a machine task at the decoder site is determined and redundant layers can be removed. Channels in at least one feature layer can be evaluated for redundancy and redundant channels also removed prior to encoding.
Resumen de: AU2025200702A1
The present invention relates to Structured Intelligence Refinement (SIR), a novel framework designed to stabilize, optimize, and regulate artificial intelligence (AI) self-improvement through controlled recursive learning cycles. This system prevents intelligence drift, over-optimization, cognitive fragmentation, and instability that commonly arise in self-modifying AI architectures. The SIR framework incorporates four core stabilization mechanisms: Recursive Intelligence Stabilization (RIS) – A multi-tiered reinforcement structure that prevents runaway recursion, ensuring AI refinements remain incremental, stable, and bounded. RIS dynamically regulates recursive depth by evaluating learning stability, performance gains, and entropy control, enforcing adaptive rollback mechanisms when instability is detected. AI Identity Core (AIC) – A persistent cognitive self-modeling framework that ensures AI retains coherence and logical consistency across recursive learning iterations. AIC prevents cognitive fragmentation by maintaining a hierarchical memory structure that tracks intelligence state changes, self-referencing prior decision pathways to ensure stable refinements. Adaptive Refinement Thresholds (ART) – A dynamic intelligence expansion regulator that modulates the frequency, magnitude, and depth of self-improvement cycles based on system confidence scores, historical stability, and human-aligned interpretability metrics. ART balances exploration vs. exploitation, ensur
Resumen de: WO2026172133A1
This invention describes an Al-supported method for dynamic decision optimization, autonomous bias correction, and scalable self-optimization. The method is based on an interdependent architecture consisting of three core modules: ⃰ Impulse Reflection Dynamics (IRD) - Identification & correction of cognitive biases using neural networks and self-learning algorithms. X Quantum Shift Discourse (QSD) - Simulation and generation of alternative decision models to optimize decision pathways. X Iterative Quantum Reflection (IQR-180°) - Real-time synchronization of reflection & action to continuously improve decision-making processes. The method is structured within a closed data flow model, ensuring that no module operates in isolation. It guarantees: √ Prevention of decision errors through self-adaptive optimization. & Mathematically defined interdependencies, preventing fragmented or modified use. √ Dynamic scaling through internal and external impulse generators, enabling adaptive adjustments to various decision contexts. Bl Application in corporate strategies, Al training & automated systems to enhance decision-making processes. This architecture enables continuous upscaling of decision optimization and can be integrated into both existing and newly developed Al systems.
Resumen de: US20260245547A1
Disclosed are apparatuses, systems, and techniques for implementing efficient transcription of multi-speaker speech with overlapping utterances using speaker activity detection. The techniques include processing, using a first set of neural network (NN) layers of an automatic speech recognition (ASR) model, speech data for the multi-speaker speech to generate an intermediate feature (IF) representative of the speech data. The techniques further include modifying, using a second set of NN layers of the ASR model, the IF to obtain modified IFs using speaker activity data, which identifies times when various speakers speak in the multi-speaker speech. The techniques further include processing the modified IFs to obtain a plurality of transcriptions identifying content of speech of the plurality of speakers, and generating, using the plurality of transcriptions, a transcript of the multi-speaker speech.
Resumen de: US20260244898A1
Provided in the present disclosure are a data processing method and apparatus, device, and medium. The method includes: inputting data to be processed into a target neural network for processing to obtain a processing result of the data to be processed, at least one convolution layer of the target neural network being an attention convolution layer based on a first attention mechanism, and/or, performing feature fusion between at least two levels of convolution layers of the target neural network on the basis of a second attention mechanism, the first attention mechanism including a self-attention mechanism for a local area of a feature, and the second attention mechanism including an attention mechanism for a local area of an output feature between output features of different scales.
Resumen de: US20260245340A1
0000 In various examples, live perception from sensors of a vehicle may be leveraged to detect and classify intersection contention areas in an environment of a vehicle in real-time or near real-time. For example, a deep neural network (DNN) may be trained to compute outputs—such as signed distance functions—that may correspond to locations of boundaries delineating intersection contention areas. The signed distance functions may be decoded and/or post-processed to determine instance segmentation masks representing locations and classifications of intersection areas or regions. The locations of the intersections areas or regions may be generated in image-space and converted to world-space coordinates to aid an autonomous or semi-autonomous vehicle in navigating intersections according to rules of the road, traffic priority considerations, and/or the like.
Resumen de: US20260245346A1
0000 One embodiment of a method includes calculating one or more activation values of one or more neural networks trained to infer eye gaze information based, at least in part, on eye position of one or more images of one or more faces indicated by an infrared light reflection from the one or more images.
Resumen de: US20260245403A1
0000 This disclosure relates generally to a method and system for predicting animal emotions using animal emotion knowledge graph and graph neural network. Current available methods focus only on visual and language data captured from animals, and lacks real time adaptability. The method disclosed generates an animal emotion knowledge graph (AEKG) that combines human and animal neurobiological data, behavioral studies, and the human wheel of emotions. Further real time graphs are generated from multimodal input data captured from the animal. These real time graphs are used for predicting primary, secondary, and tertiary emotions of the animals using a trained graph neural network-Transformer model. This model is trained using the AEKG. Using temporal graph analysis, the method predicts future emotions and generates real-time recommendations based on generative artificial intelligence techniques. Predicting the emotions of animals in real time helps to grasp their emotional well-being to improve their care and management effectively.
Resumen de: US20260244926A1
Embodiments described herein provide a method for hardware resource allocation during the operation of a generative neural network model. The method includes receiving a set of input data at a neural network-based model implemented on one or more hardware processors, where the model comprises a plurality of sequentially connected blocks. The method involves computing respective input and output intermediate values for at least one block during forward passes of the model, and calculating a change in entropy for each block based on the difference between entropy estimates for the input and output intermediate values. Blocks are pruned based on their respective changes in entropy, and hardware resources allocated to the pruned model are adjusted accordingly. The pruned neural network model is then operated using the adjusted hardware resources.
Resumen de: US20260245223A1
0000 A system receives 2D depth-map images of an empty conveyor belt. The system generates a training set of labelled 2D depth-map images based on the 2D depth-map images. The system trains or fine-tunes a neural network using the training set to form trained parameters that cause the neural network to detect items on the conveyor belt using 2D depth-map images of the items on the conveyor belt. In an alternative embodiment, a pre-trained neural network is configured to select feature map channels generated from 2D depth-map images of items on a conveyor belt. The selected channels are resized and combined to provide relevant data for generating a segmentation mask based of the items using a binarization threshold that is automatically determined. The system may include a pretrained image segmentation model that generates at least one instance segmentation mask from the 2D depth-map images.
Resumen de: EP4793824A2
0001 A hardware circuit for implementing a neural network comprising a plurality of neural network layers comprises a controller. The controller is configured to analyze output activations computed by a first compute system for a first neural network layer, where the output activations are provided on an output activation bus. The controller is further configured to determine which of the output activations have a non-zero value, generate an additional representation of the output activations that identifies only the output activations having a non-zero value, and use the additional representation to supply only the output activations having a non-zero value as input activations to a subsequent, second compute system for a second neural network layer.
Resumen de: EP4793903A2
A method of controlling an electronic apparatus includes acquiring an image and depth information of the acquired image; inputting the acquired image into a neural network model trained to acquire information on objects included in the acquired image; acquiring an intermediate feature value output by an intermediate layer of the neural network model; identifying a feature area for at least one object among the objects included in the acquired image based on the intermediate feature value; and acquiring distance information between the electronic apparatus and the at least one object based on the feature area for the at least one object and the depth information.
Resumen de: EP4793878A1
Provided are an image processing device and an operating method of the same. The image processing device includes a memory storing one or more instructions, and at least one processor including processing circuitry, and memory storing one or more instructions that, when executed by the at least one processor individually or collectively, cause the image processing device to obtain a neural network model corresponding to a quality of an input image and viewing information related to the input image. The at least one processor is configured to generate training data, based on the quality of the input image and the viewing information. The at least one processor is configured to train the neural network model by using the training data. The at least one processor is configured to obtain an image quality processed output image from the input image, based on the trained neural network model.
Resumen de: EP4793797A1
The present disclosure relates to the field of control, and provides a training method and apparatus, a vehicle safety function control method and apparatus, and a vehicle. The training method comprises : acquiring a signal sample image of a vehicle, wherein the signal sample image comprises a safety function normal image and a safety function failure image; using the signal sample image to train a deep neural network model, and using the trained deep neural network model to perform data augmentation processing on the signal sample image to obtain training sample images; and using the training sample images to train a safety function failure identification classifier, wherein the safety function failure identification classifier is used for identifying a safety function state during vehicle operation, and the safety function state includes a safety function normal state or a safety function failure state.
Resumen de: EP4793799A1
The present application provides a data processing method and apparatus based on multimodal fusion, pertaining to the technical field of data processing, where the method includes: acquiring one-dimensional data and image data; converting the one-dimensional data into two-dimensional data based on a dimension of the image data; performing zero-padding processing on vacant positions in the two-dimensional data; performing stacking processing on the zero-padded two-dimensional data and the image data to obtain a multilayer stacked input feature map; performing fusion processing on the multilayer stacked input feature map through a neural network to obtain a fused feature map; and performing data processing based on the fused feature map. The present invention can unify the data formats of different modalities, enabling them to be processed in the same feature space, significantly simplifying the alignment process between heterogeneous data.
Resumen de: EP4793908A1
0001 Disclosed is a computer-implemented method for profiling particles in a taxa using a sequence of input data images. The process involves identifying and categorizing suspended particles in each image to obtain bounding boxes and classification data. These particles are then tracked across subsequent images using the bounding boxes to compile tracking data. The method uses this data to output a profile of the suspended particles, incorporating taxonomic identification and possibly using convolutional operations. It employs two neural networks: one for identifying regions of interest and another for categorizing the particles based on these regions. The profile may include biomass calculations and assessments of ecosystem status, integrating sensor metadata such as depth, chlorophyll-a, salinity, and temperature. Non-particle elements like bubbles and damaged areas are excluded from tracking. The method also encompasses a system setup with a camera and processor, and a computer program that enables the execution of these methods.
Resumen de: US20260237189A1
0000 An apparatus for processing image data includes a memory for storing the image data and processing circuitry in communication with the memory. The processing circuitry is configured to obtain image data including a current set of multiple camera images from multiple cameras. According to such an example, the apparatus may also generate respective feature vectors from each of the multiple camera images with a shared image feature encoder using camera-specific positional embeddings associated with different respective cameras used to capture the multiple camera images. The apparatus may also perform a perception task using the respective feature vectors.
Resumen de: WO2026169639A1
A system for training at least one graph neural network (GNN) for monitoring organ health includes at least one memory configured to store instructions and at least one processor configured to execute the instructions to cause the system to obtain a plurality of parameters from healthy individuals and diseased individuals, perform causal discovery to determine one or more causal relationships between the plurality of organ parameters, train the at least one GNN with the one or more causal relationships to create a trained GNN configured to output a multi-organ health status based on patient parameters. The system is also configured to train at least one acute event prediction model to predict at least one acute or adverse event for at least one organ
Resumen de: WO2026169963A1
A method for optimizing device settings of a plurality of devices includes receiving a performance characteristics target for devices manufactured according to a common design, wherein each device is configurable via device settings and exhibits performance variations over process variations and device operating conditions. The method includes defining a plurality of test cases, each including a combination of a respective device, a process variation, and a device operating condition. The method includes applying a first genetic adaptation algorithm to produce test results, training a neural network to generate predicted performance characteristics, and determining, by using the trained neural network as a surrogate model and applying a second genetic adaptation algorithm, a global set of device settings that collectively achieves the performance characteristics target across the plurality of devices.
Resumen de: US20260236740A1
Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for controlling an agent interacting with an environment using a Transformer neural network.
Resumen de: WO2026170170A1
A method for providing a Generative Artificial Intelligence (GenAI) system with trustworthiness evaluation, wherein the method comprises processing a training dataset that the GenAI was trained upon; constructing a plurality of latent knowledge anchors (LKAs) by applying topic modeling on the processed training dataset, wherein the plurality of LKAs comprises domain-specific knowledge representations within the processed training dataset; training a neural network classifier using the LKAs and the processed training data; evaluating, using the neural network classifier, a response generated by the GenAI to generate a trust score, wherein the trust score indicates the alignment level of the response to the training dataset and/or the comprehensiveness of the response compared to the training dataset; in response to the trust score exceeding a threshold, presenting the response to a user; and in response to the trust score not exceeding the threshold, issuing a command for the GenAI to conduct additional actions.
Resumen de: WO2026169910A1
Systems, apparatuses, and methods disclosed herein relate to a digital image processing system. In one aspect, the digital image processing system includes one or more processors and one or more computer-readable media storing instructions which, when executed by the one or more processors, cause the digital image processing system to access a digital image depicting a portion of a biological sample obtained from a subject. A projecting space and a staining category for an image conversion is defined by the digital image processing system, causing the digital image to be projected to the projecting space using a neural network model. The projected image is output by the digital image processing system.
Nº publicación: US20260237074A1 13/08/2026
Solicitante:
NVIDIA CORP [US]
NVIDIA Corporation
Resumen de: US20260237074A1
Class agnostic object mask generation uses a vision transformer-based auto-labeling framework requiring only images and object bounding boxes to generate object (segmentation) masks. The generated object masks, images, and object labels may then be used to train instance segmentation models or other neural networks to localize and segment objects with pixel-level accuracy.