Resumen de: US20260244913A1
A block-centric acceleration system and method for heterogeneous graph neural networks (HGNNs) are provided. The system includes a processor comprising a block loading unit, a block scheduling unit, processing units, and a reduction unit. The block loading unit identifies central blocks and constructs a block overlap graph, prioritizing the central blocks based on their degrees of overlap. The block scheduling unit dynamically assigns prioritized blocks to idle processing units. The processing units perform computations using row-wise matrix multiplication and element-wise operations, thereby carrying out hybrid computation for HGNN inference and generating intermediate results corresponding to structural and semantic aggregation within the block overlap graph. The reduction unit determines an aggregation status flag based on results from the processing units and outputs semantic features upon completion of structural and semantic aggregation. The present disclosure reduces redundant feature access, improves processing efficiency, and lowers memory bandwidth requirements.
Resumen de: WO2026170844A1
Embodiments of the present disclosure relate to the technical field of artificial intelligence, and provide a data processing method, a device, a medium, and a program product. The data processing method comprises: performing multi-modal encoding processing on a target question to obtain multi-modal question encoding, and determining multi-modal data corresponding to the target question; on the basis of the multi-modal question encoding, determining target reference data corresponding to the target question from among the multi-modal data, wherein the target reference data is at least one modality of data among the multi-modal data; and using a question processing model to process the target question on the basis of the target reference data, to obtain a question processing result corresponding to the target question. The problem of inaccurate data processing results caused by limited knowledge learned by neural network models is avoided.
Resumen de: US20260244926A1
Embodiments described herein provide a method for hardware resource allocation during the operation of a generative neural network model. The method includes receiving a set of input data at a neural network-based model implemented on one or more hardware processors, where the model comprises a plurality of sequentially connected blocks. The method involves computing respective input and output intermediate values for at least one block during forward passes of the model, and calculating a change in entropy for each block based on the difference between entropy estimates for the input and output intermediate values. Blocks are pruned based on their respective changes in entropy, and hardware resources allocated to the pruned model are adjusted accordingly. The pruned neural network model is then operated using the adjusted hardware resources.
Resumen de: US20260245041A1
Disclosed is a graph neural network (GNN) and transformer-based task planner. A task planning method performed by a task planning system may include constructing a task planner model based on a graph neural network and a transformer; and generating a task plan from a goal instruction of a task and history information of a scene graph through the constructed task planner model.
Resumen de: US20260245403A1
0000 This disclosure relates generally to a method and system for predicting animal emotions using animal emotion knowledge graph and graph neural network. Current available methods focus only on visual and language data captured from animals, and lacks real time adaptability. The method disclosed generates an animal emotion knowledge graph (AEKG) that combines human and animal neurobiological data, behavioral studies, and the human wheel of emotions. Further real time graphs are generated from multimodal input data captured from the animal. These real time graphs are used for predicting primary, secondary, and tertiary emotions of the animals using a trained graph neural network-Transformer model. This model is trained using the AEKG. Using temporal graph analysis, the method predicts future emotions and generates real-time recommendations based on generative artificial intelligence techniques. Predicting the emotions of animals in real time helps to grasp their emotional well-being to improve their care and management effectively.
Resumen de: US20260244977A1
0000 The concordance based artificial intelligence model utilizes human or program defined keywords to analyze data and create output based on its analysis. This analysis is comprised of layers of program generated concordances and concatenations of concordances to produce content from its source data that is contextually relevant to queries. The concordance based artificial intelligence model is immune from the “hallucinations” phenomenon found in neural network based artificial intelligence models. It mimics human memory in its operation by producing a full forensic path of each of its operations in the form of text files that are saved for future use and are used to train the model over time.
Resumen de: US20260245547A1
Disclosed are apparatuses, systems, and techniques for implementing efficient transcription of multi-speaker speech with overlapping utterances using speaker activity detection. The techniques include processing, using a first set of neural network (NN) layers of an automatic speech recognition (ASR) model, speech data for the multi-speaker speech to generate an intermediate feature (IF) representative of the speech data. The techniques further include modifying, using a second set of NN layers of the ASR model, the IF to obtain modified IFs using speaker activity data, which identifies times when various speakers speak in the multi-speaker speech. The techniques further include processing the modified IFs to obtain a plurality of transcriptions identifying content of speech of the plurality of speakers, and generating, using the plurality of transcriptions, a transcript of the multi-speaker speech.
Resumen de: US20260241941A1
A computer-implemented method for monitoring an artificial deep neural network comprises supplying input data to the trained deep neural network monitored, in order to obtain therefrom output data activation map data, and supplying the input data, the output data and the activation map data to a computer-implemented network observer. The network observer generates masking data from the activation map data and/or the input data and/or the output data; masks the activation map data using the masking data in order to obtain masked activation map data, wherein the masked activation map data contain unmasked activation values and masked activation values; and determines an outlier score for the output data using the masked activation map data, wherein merely the unmasked values are taken into account when determining the outlier score. The outlier score is a numerical value and indicates the extent to which the determined output deviates from a typical case.
Resumen de: US20260244939A1
0000 According to one aspect of the embodiments, a method includes acquiring vital data, behavior data, and task data over a predetermined period of users to be learned from a learning data storage unit and performing predetermined preprocessing on each of the acquired vital data, the acquired behavior data, and the acquired task data. The method includes generating feature data regarding the users to be learned by combining the preprocessed vital data, the preprocessed behavior data, and the preprocessed task data of the same users to be learned while aligning time-series positions. The method includes generating a learned analysis model by learning an analysis model having a neural network structure constructed in advance using the feature data and present bias data regarding the users to be learned.
Resumen de: US20260245223A1
0000 A system receives 2D depth-map images of an empty conveyor belt. The system generates a training set of labelled 2D depth-map images based on the 2D depth-map images. The system trains or fine-tunes a neural network using the training set to form trained parameters that cause the neural network to detect items on the conveyor belt using 2D depth-map images of the items on the conveyor belt. In an alternative embodiment, a pre-trained neural network is configured to select feature map channels generated from 2D depth-map images of items on a conveyor belt. The selected channels are resized and combined to provide relevant data for generating a segmentation mask based of the items using a binarization threshold that is automatically determined. The system may include a pretrained image segmentation model that generates at least one instance segmentation mask from the 2D depth-map images.
Resumen de: WO2026170664A1
The present invention relates to the technical field of long-tailed medical image classification. Disclosed are a contrastive-learning-based medical image classification method and system, and a storage medium. The method comprises: dividing a long-tailed medical image data set into a training set and a test set, and on the basis of a preset scheme and in the form of batch processing, separately performing weak data augmentation and strong data augmentation on images in the training set; by means of a deep convolutional neural network, performing a contrastive learning task on obtained weakly data-augmented images and obtained strongly data-augmented images, and learning network parameters to obtain a parameter-optimized deep convolutional neural network; and classifying long-tailed medical images in the test set. In the present application, by means of a prototype-enhanced contrastive learning strategy, learnable-class prototypes are generated and subjected to data augmentation, so as to obtain a balanced implicitly-augmented contrastive learning loss, thereby achieving the advantages of high precision, high efficiency, low costs, wide applicability, etc.
Resumen de: US20260246928A1
0000 Systems and methods are provided for encoding and decoding video for machine consumption in which bandwidth is reduced by filtering feature layers and filtering channels at the encoder site that are determined to be redundant or of reduced relevance. A video encoder includes a neural network front end which receives image data and generates a plurality of feature layers. The relevance of the plurality of feature layers to a machine task at the decoder site is determined and redundant layers can be removed. Channels in at least one feature layer can be evaluated for redundancy and redundant channels also removed prior to encoding.
Resumen de: US20260244680A1
Embodiments described herein provide a method of arithmetic reasoning generation by an artificial intelligence (AI) agent. The method includes generating, an image containing at least one object having a target arithmetic property; generating a query relating to the target arithmetic property; generating, by a first neural network language model, a positive response and a negative response; and forming a training quadruple including the image, the query, the positive response and the negative response. The method also includes generating, by a second neural network language model a candidate response associated with a first probability that the candidate response is the positive response and a second probability that the candidate response is the negative response, training the second neural network language model; and building, at a server the AI agent employing the second neural network based language model after the training.
Resumen de: EP4793878A1
Provided are an image processing device and an operating method of the same. The image processing device includes a memory storing one or more instructions, and at least one processor including processing circuitry, and memory storing one or more instructions that, when executed by the at least one processor individually or collectively, cause the image processing device to obtain a neural network model corresponding to a quality of an input image and viewing information related to the input image. The at least one processor is configured to generate training data, based on the quality of the input image and the viewing information. The at least one processor is configured to train the neural network model by using the training data. The at least one processor is configured to obtain an image quality processed output image from the input image, based on the trained neural network model.
Resumen de: EP4793903A2
A method of controlling an electronic apparatus includes acquiring an image and depth information of the acquired image; inputting the acquired image into a neural network model trained to acquire information on objects included in the acquired image; acquiring an intermediate feature value output by an intermediate layer of the neural network model; identifying a feature area for at least one object among the objects included in the acquired image based on the intermediate feature value; and acquiring distance information between the electronic apparatus and the at least one object based on the feature area for the at least one object and the depth information.
Resumen de: EP4793824A2
0001 A hardware circuit for implementing a neural network comprising a plurality of neural network layers comprises a controller. The controller is configured to analyze output activations computed by a first compute system for a first neural network layer, where the output activations are provided on an output activation bus. The controller is further configured to determine which of the output activations have a non-zero value, generate an additional representation of the output activations that identifies only the output activations having a non-zero value, and use the additional representation to supply only the output activations having a non-zero value as input activations to a subsequent, second compute system for a second neural network layer.
Resumen de: EP4793797A1
The present disclosure relates to the field of control, and provides a training method and apparatus, a vehicle safety function control method and apparatus, and a vehicle. The training method comprises : acquiring a signal sample image of a vehicle, wherein the signal sample image comprises a safety function normal image and a safety function failure image; using the signal sample image to train a deep neural network model, and using the trained deep neural network model to perform data augmentation processing on the signal sample image to obtain training sample images; and using the training sample images to train a safety function failure identification classifier, wherein the safety function failure identification classifier is used for identifying a safety function state during vehicle operation, and the safety function state includes a safety function normal state or a safety function failure state.
Resumen de: EP4793799A1
The present application provides a data processing method and apparatus based on multimodal fusion, pertaining to the technical field of data processing, where the method includes: acquiring one-dimensional data and image data; converting the one-dimensional data into two-dimensional data based on a dimension of the image data; performing zero-padding processing on vacant positions in the two-dimensional data; performing stacking processing on the zero-padded two-dimensional data and the image data to obtain a multilayer stacked input feature map; performing fusion processing on the multilayer stacked input feature map through a neural network to obtain a fused feature map; and performing data processing based on the fused feature map. The present invention can unify the data formats of different modalities, enabling them to be processed in the same feature space, significantly simplifying the alignment process between heterogeneous data.
Resumen de: EP4793908A1
0001 Disclosed is a computer-implemented method for profiling particles in a taxa using a sequence of input data images. The process involves identifying and categorizing suspended particles in each image to obtain bounding boxes and classification data. These particles are then tracked across subsequent images using the bounding boxes to compile tracking data. The method uses this data to output a profile of the suspended particles, incorporating taxonomic identification and possibly using convolutional operations. It employs two neural networks: one for identifying regions of interest and another for categorizing the particles based on these regions. The profile may include biomass calculations and assessments of ecosystem status, integrating sensor metadata such as depth, chlorophyll-a, salinity, and temperature. Non-particle elements like bubbles and damaged areas are excluded from tracking. The method also encompasses a system setup with a camera and processor, and a computer program that enables the execution of these methods.
Resumen de: WO2026170170A1
A method for providing a Generative Artificial Intelligence (GenAI) system with trustworthiness evaluation, wherein the method comprises processing a training dataset that the GenAI was trained upon; constructing a plurality of latent knowledge anchors (LKAs) by applying topic modeling on the processed training dataset, wherein the plurality of LKAs comprises domain-specific knowledge representations within the processed training dataset; training a neural network classifier using the LKAs and the processed training data; evaluating, using the neural network classifier, a response generated by the GenAI to generate a trust score, wherein the trust score indicates the alignment level of the response to the training dataset and/or the comprehensiveness of the response compared to the training dataset; in response to the trust score exceeding a threshold, presenting the response to a user; and in response to the trust score not exceeding the threshold, issuing a command for the GenAI to conduct additional actions.
Resumen de: WO2026169535A1
Apparatuses, systems, and techniques to identify information to evict from a Key-Value (KV) cache. In at least one embodiment, information stored within one or more large language model (LLM) KV caches may be identified to cause an indication to be generated of information stored within the one or more LLM KV caches that can be evicted without causing other information stored within the one or more LLM KV caches to be restored to the one or more LLM KV caches.
Resumen de: AU2025214687A1
Disclosed herein are sequencing systems and sequencing methods for training neural networks and for utilizing the trained neural networks for sequencing analysis after acquiring flow cell images using the sequencing systems. The sequencing systems disclosed herein can include Field-Programmable Gate Array (FPGAs), artificial intelligence (AI) chips, or a combination thereof.
Resumen de: WO2026165976A1
The present invention relates to the technical field of AI, and provides an AI-based method and system for automatic optimization of distributed computing tasks for big data. The method comprises: constructing a multi-modal spatiotemporal feature perception network to extract spatiotemporal features of tasks and resources, and training a hierarchical hybrid decision network to perform global task scheduling and local resource allocation. A hierarchical deep reinforcement learning architecture is used for a global policy network, which combines Monte Carlo tree search and prioritized experience replay to generate decisions. A local execution network optimizes local deployment on the basis of a graph neural network and multi-agent collaborative learning. In addition, a distributed anomaly detection network is deployed to monitor performance in real time, adaptive tuning is achieved by means of reinforcement transfer learning, and finally the perception network is updated by means of knowledge distillation, thereby achieving model evolution. The present invention can effectively improve the execution efficiency and resource utilization rate of distributed computing tasks for big data, and reduce system operation costs.
Resumen de: US20260237377A1
0000 There is provided a method for synthesizing a speech waveform from text data. The method comprises: determining, from the text data, a phoneme sequence; obtaining a reference speech waveform comprising a high-level speech representation of a reference speaker speaking a reference speech; applying a trained neural network to the high-level speech representation to extract speaker embeddings; and determining, from the speaker embeddings and the phoneme sequence, a synthesized speech waveform indicative of the reference speaker speaking the text data.
Nº publicación: US20260237019A1 13/08/2026
Solicitante:
GOOGLE LLC [US]
Google LLC
Resumen de: US20260237019A1
0000 Methods, systems, and apparatus, including computer programs encoded on computer storage media, for generating images of a new subject using a diffusion neural network.