Resumen de: EP4797161A2
A system and method of calibrating a broadcast video feed are disclosed herein. A computing system retrieves a plurality of broadcast video feeds that include a plurality of video frames. The computing system generates a trained neural network, by generating a plurality of training data sets based on the broadcast video feed and learning, by the neural network, to generate a homography matrix for each frame of the plurality of frames. The computing system receives a target broadcast video feed for a target sporting event. The computing system partitions the target broadcast video feed into a plurality of target frames. The computing system generates for each target frame in the plurality of target frames, via the neural network, a target homography matrix. The computing system calibrates the target broadcast video feed by warping each target frame by a respective target homography matrix.
Resumen de: EP4797714A1
An electronic device for performing vision perception from an image acquired using a meta lens, and an operating method thereof are provided. The electronic device according to one embodiment of the present disclosure comprises: a meta lens having a pattern formed on the surface thereof and including a plurality of pillars or pins with different shapes, heights and widths; and an image sensor configured to receive phase-modulated light reflected from an object and transmitted through the meta lens, and obtain a coded image by converting the received light into an electrical signal; and at least one processor configured to input the coded image into an artificial intelligence model, and obtain a label indicating a perception result of an object through inference using the artificial intelligence model, wherein the artificial intelligence model may be a neural network model trained to obtain a simulation image by inputting an RGB image into a model reflecting optical characteristics of the meta lens, and output, as the perception result of the simulation image, a label indicating ground truth of the RGB image that was input.
Resumen de: EP4797152A1
The invention relates to a computer-implemented method for approximating at least two unknown variables (10) of a set of partial differential equations using a quantum physics-informed neural network (100), the method comprising the steps: providing a numerical grid (20) having grid coordinates (30) at which the unknown variables (10) are to be approximated; providing a quantum physics-informed neural network (100), the quantum physics-informed neural network (100) comprising a hybrid network (110) for each unknown variable (10), wherein the hybrid networks (110) are not interconnected to each other, each hybrid network (110) having a quantum network (120) having at least one quantum layer (130) and a classical network (140) having at least one classical layer (150), wherein the respective quantum network (120) and the respective classical network (140) are not interconnected; inputting the grid coordinates (30) into the quantum physics-informed neural network (100), such that the grid coordinates (30) are input to each hybrid network (110); and computing the output (Hout) of each hybrid network (110) for each grid coordinate (30), each output (Hout) corresponding to a different unknown variable (10) to be approximated at the grid coordinate (30), wherein the output (Hout) of each hybrid network (110) is obtained by combining an output of the respective quantum network (Qout) and an output of the respective classical network (C
Resumen de: WO2025109032A2
Methods, systems, and apparatus, including computer programs encoded on computer storage media, for generating audio and, optionally, a corresponding image using generative neural networks. For example, a spectrogram of the audio can be generated using a hierarchy of diffusion neural networks.
Resumen de: US20260245340A1
0000 In various examples, live perception from sensors of a vehicle may be leveraged to detect and classify intersection contention areas in an environment of a vehicle in real-time or near real-time. For example, a deep neural network (DNN) may be trained to compute outputs—such as signed distance functions—that may correspond to locations of boundaries delineating intersection contention areas. The signed distance functions may be decoded and/or post-processed to determine instance segmentation masks representing locations and classifications of intersection areas or regions. The locations of the intersections areas or regions may be generated in image-space and converted to world-space coordinates to aid an autonomous or semi-autonomous vehicle in navigating intersections according to rules of the road, traffic priority considerations, and/or the like.
Resumen de: US20260245346A1
0000 One embodiment of a method includes calculating one or more activation values of one or more neural networks trained to infer eye gaze information based, at least in part, on eye position of one or more images of one or more faces indicated by an infrared light reflection from the one or more images.
Resumen de: US20260244898A1
Provided in the present disclosure are a data processing method and apparatus, device, and medium. The method includes: inputting data to be processed into a target neural network for processing to obtain a processing result of the data to be processed, at least one convolution layer of the target neural network being an attention convolution layer based on a first attention mechanism, and/or, performing feature fusion between at least two levels of convolution layers of the target neural network on the basis of a second attention mechanism, the first attention mechanism including a self-attention mechanism for a local area of a feature, and the second attention mechanism including an attention mechanism for a local area of an output feature between output features of different scales.
Resumen de: US20260245403A1
0000 This disclosure relates generally to a method and system for predicting animal emotions using animal emotion knowledge graph and graph neural network. Current available methods focus only on visual and language data captured from animals, and lacks real time adaptability. The method disclosed generates an animal emotion knowledge graph (AEKG) that combines human and animal neurobiological data, behavioral studies, and the human wheel of emotions. Further real time graphs are generated from multimodal input data captured from the animal. These real time graphs are used for predicting primary, secondary, and tertiary emotions of the animals using a trained graph neural network-Transformer model. This model is trained using the AEKG. Using temporal graph analysis, the method predicts future emotions and generates real-time recommendations based on generative artificial intelligence techniques. Predicting the emotions of animals in real time helps to grasp their emotional well-being to improve their care and management effectively.
Resumen de: US20260244928A1
0000 A system for interfacing a persistent cognitive machine with a legacy neural network through PCM-enhanced supervisory neurons that operate as bidirectional projection interfaces. Each supervisory neuron maps temporal sequences of activation states from a monitored local neural network region into trajectories within a cognitive configuration space, where geometric analysis computes curvature estimates, holonomy signatures, boundary mismatch functionals, and homotopy class identifications. A holonomy accumulator stores compressed holonomy representations that grow logarithmically with accumulated experience. A variational modification planner selects structural modifications by computing stationary trajectories of an action functional encoding cost, holonomy-derived bias, and boundary mismatch penalties. Failed modifications are irreversibly exported to a residual sector whose curvature structure prevents gradient return, ensuring monotonic improvement. A persistent cognitive substrate maintains accumulated holonomy and residual constraints across inference sessions, enabling the system to continuously adapt the legacy neural network without repeating known harmful strategies.
Resumen de: US20260244680A1
Embodiments described herein provide a method of arithmetic reasoning generation by an artificial intelligence (AI) agent. The method includes generating, an image containing at least one object having a target arithmetic property; generating a query relating to the target arithmetic property; generating, by a first neural network language model, a positive response and a negative response; and forming a training quadruple including the image, the query, the positive response and the negative response. The method also includes generating, by a second neural network language model a candidate response associated with a first probability that the candidate response is the positive response and a second probability that the candidate response is the negative response, training the second neural network language model; and building, at a server the AI agent employing the second neural network based language model after the training.
Resumen de: US20260245223A1
0000 A system receives 2D depth-map images of an empty conveyor belt. The system generates a training set of labelled 2D depth-map images based on the 2D depth-map images. The system trains or fine-tunes a neural network using the training set to form trained parameters that cause the neural network to detect items on the conveyor belt using 2D depth-map images of the items on the conveyor belt. In an alternative embodiment, a pre-trained neural network is configured to select feature map channels generated from 2D depth-map images of items on a conveyor belt. The selected channels are resized and combined to provide relevant data for generating a segmentation mask based of the items using a binarization threshold that is automatically determined. The system may include a pretrained image segmentation model that generates at least one instance segmentation mask from the 2D depth-map images.
Resumen de: US20260245547A1
Disclosed are apparatuses, systems, and techniques for implementing efficient transcription of multi-speaker speech with overlapping utterances using speaker activity detection. The techniques include processing, using a first set of neural network (NN) layers of an automatic speech recognition (ASR) model, speech data for the multi-speaker speech to generate an intermediate feature (IF) representative of the speech data. The techniques further include modifying, using a second set of NN layers of the ASR model, the IF to obtain modified IFs using speaker activity data, which identifies times when various speakers speak in the multi-speaker speech. The techniques further include processing the modified IFs to obtain a plurality of transcriptions identifying content of speech of the plurality of speakers, and generating, using the plurality of transcriptions, a transcript of the multi-speaker speech.
Resumen de: US20260244977A1
0000 The concordance based artificial intelligence model utilizes human or program defined keywords to analyze data and create output based on its analysis. This analysis is comprised of layers of program generated concordances and concatenations of concordances to produce content from its source data that is contextually relevant to queries. The concordance based artificial intelligence model is immune from the “hallucinations” phenomenon found in neural network based artificial intelligence models. It mimics human memory in its operation by producing a full forensic path of each of its operations in the form of text files that are saved for future use and are used to train the model over time.
Resumen de: WO2026172133A1
This invention describes an Al-supported method for dynamic decision optimization, autonomous bias correction, and scalable self-optimization. The method is based on an interdependent architecture consisting of three core modules: ⃰ Impulse Reflection Dynamics (IRD) - Identification & correction of cognitive biases using neural networks and self-learning algorithms. X Quantum Shift Discourse (QSD) - Simulation and generation of alternative decision models to optimize decision pathways. X Iterative Quantum Reflection (IQR-180°) - Real-time synchronization of reflection & action to continuously improve decision-making processes. The method is structured within a closed data flow model, ensuring that no module operates in isolation. It guarantees: √ Prevention of decision errors through self-adaptive optimization. & Mathematically defined interdependencies, preventing fragmented or modified use. √ Dynamic scaling through internal and external impulse generators, enabling adaptive adjustments to various decision contexts. Bl Application in corporate strategies, Al training & automated systems to enhance decision-making processes. This architecture enables continuous upscaling of decision optimization and can be integrated into both existing and newly developed Al systems.
Resumen de: AU2025200702A1
The present invention relates to Structured Intelligence Refinement (SIR), a novel framework designed to stabilize, optimize, and regulate artificial intelligence (AI) self-improvement through controlled recursive learning cycles. This system prevents intelligence drift, over-optimization, cognitive fragmentation, and instability that commonly arise in self-modifying AI architectures. The SIR framework incorporates four core stabilization mechanisms: Recursive Intelligence Stabilization (RIS) – A multi-tiered reinforcement structure that prevents runaway recursion, ensuring AI refinements remain incremental, stable, and bounded. RIS dynamically regulates recursive depth by evaluating learning stability, performance gains, and entropy control, enforcing adaptive rollback mechanisms when instability is detected. AI Identity Core (AIC) – A persistent cognitive self-modeling framework that ensures AI retains coherence and logical consistency across recursive learning iterations. AIC prevents cognitive fragmentation by maintaining a hierarchical memory structure that tracks intelligence state changes, self-referencing prior decision pathways to ensure stable refinements. Adaptive Refinement Thresholds (ART) – A dynamic intelligence expansion regulator that modulates the frequency, magnitude, and depth of self-improvement cycles based on system confidence scores, historical stability, and human-aligned interpretability metrics. ART balances exploration vs. exploitation, ensur
Resumen de: WO2026170844A1
Embodiments of the present disclosure relate to the technical field of artificial intelligence, and provide a data processing method, a device, a medium, and a program product. The data processing method comprises: performing multi-modal encoding processing on a target question to obtain multi-modal question encoding, and determining multi-modal data corresponding to the target question; on the basis of the multi-modal question encoding, determining target reference data corresponding to the target question from among the multi-modal data, wherein the target reference data is at least one modality of data among the multi-modal data; and using a question processing model to process the target question on the basis of the target reference data, to obtain a question processing result corresponding to the target question. The problem of inaccurate data processing results caused by limited knowledge learned by neural network models is avoided.
Resumen de: WO2026173773A1
An apparatus for processing image data includes a memory for storing the image data and processing circuitry in communication with the memory. The processing circuitry is configured to obtain image data including a current set of multiple camera images from multiple cameras. According to such an example, the apparatus may also generate respective feature vectors from each of the multiple camera images with a shared image feature encoder using camera-specific positional embeddings associated with different respective cameras used to capture the multiple camera images. The apparatus may also perform a perception task using the respective feature vectors.
Resumen de: US20260244971A1
This disclosure proposes a training method for quantum machine learning and an electronic device. The training method includes: configuring a quantum circuit to output probabilities of multiple qubits, where the quantum circuit comprises multiple gates with circuit parameters; mapping the qubits to multiple model parameters of a neural network, where multiple bases are calculated based on the qubits, and the quantity of the bases is greater than or equal to the quantity of the model parameters; inputting data into the neural network and calculating a loss based on the output of the neural network; and updating the circuit parameters in the quantum circuit according to the loss.
Resumen de: US20260244939A1
0000 According to one aspect of the embodiments, a method includes acquiring vital data, behavior data, and task data over a predetermined period of users to be learned from a learning data storage unit and performing predetermined preprocessing on each of the acquired vital data, the acquired behavior data, and the acquired task data. The method includes generating feature data regarding the users to be learned by combining the preprocessed vital data, the preprocessed behavior data, and the preprocessed task data of the same users to be learned while aligning time-series positions. The method includes generating a learned analysis model by learning an analysis model having a neural network structure constructed in advance using the feature data and present bias data regarding the users to be learned.
Resumen de: US20260241941A1
A computer-implemented method for monitoring an artificial deep neural network comprises supplying input data to the trained deep neural network monitored, in order to obtain therefrom output data activation map data, and supplying the input data, the output data and the activation map data to a computer-implemented network observer. The network observer generates masking data from the activation map data and/or the input data and/or the output data; masks the activation map data using the masking data in order to obtain masked activation map data, wherein the masked activation map data contain unmasked activation values and masked activation values; and determines an outlier score for the output data using the masked activation map data, wherein merely the unmasked values are taken into account when determining the outlier score. The outlier score is a numerical value and indicates the extent to which the determined output deviates from a typical case.
Resumen de: US20260244913A1
A block-centric acceleration system and method for heterogeneous graph neural networks (HGNNs) are provided. The system includes a processor comprising a block loading unit, a block scheduling unit, processing units, and a reduction unit. The block loading unit identifies central blocks and constructs a block overlap graph, prioritizing the central blocks based on their degrees of overlap. The block scheduling unit dynamically assigns prioritized blocks to idle processing units. The processing units perform computations using row-wise matrix multiplication and element-wise operations, thereby carrying out hybrid computation for HGNN inference and generating intermediate results corresponding to structural and semantic aggregation within the block overlap graph. The reduction unit determines an aggregation status flag based on results from the processing units and outputs semantic features upon completion of structural and semantic aggregation. The present disclosure reduces redundant feature access, improves processing efficiency, and lowers memory bandwidth requirements.
Resumen de: US20260246928A1
0000 Systems and methods are provided for encoding and decoding video for machine consumption in which bandwidth is reduced by filtering feature layers and filtering channels at the encoder site that are determined to be redundant or of reduced relevance. A video encoder includes a neural network front end which receives image data and generates a plurality of feature layers. The relevance of the plurality of feature layers to a machine task at the decoder site is determined and redundant layers can be removed. Channels in at least one feature layer can be evaluated for redundancy and redundant channels also removed prior to encoding.
Resumen de: WO2026170664A1
The present invention relates to the technical field of long-tailed medical image classification. Disclosed are a contrastive-learning-based medical image classification method and system, and a storage medium. The method comprises: dividing a long-tailed medical image data set into a training set and a test set, and on the basis of a preset scheme and in the form of batch processing, separately performing weak data augmentation and strong data augmentation on images in the training set; by means of a deep convolutional neural network, performing a contrastive learning task on obtained weakly data-augmented images and obtained strongly data-augmented images, and learning network parameters to obtain a parameter-optimized deep convolutional neural network; and classifying long-tailed medical images in the test set. In the present application, by means of a prototype-enhanced contrastive learning strategy, learnable-class prototypes are generated and subjected to data augmentation, so as to obtain a balanced implicitly-augmented contrastive learning loss, thereby achieving the advantages of high precision, high efficiency, low costs, wide applicability, etc.
Resumen de: US20260244926A1
Embodiments described herein provide a method for hardware resource allocation during the operation of a generative neural network model. The method includes receiving a set of input data at a neural network-based model implemented on one or more hardware processors, where the model comprises a plurality of sequentially connected blocks. The method involves computing respective input and output intermediate values for at least one block during forward passes of the model, and calculating a change in entropy for each block based on the difference between entropy estimates for the input and output intermediate values. Blocks are pruned based on their respective changes in entropy, and hardware resources allocated to the pruned model are adjusted accordingly. The pruned neural network model is then operated using the adjusted hardware resources.
Nº publicación: US20260245041A1 20/08/2026
Solicitante:
IUCF HYU [KR]
IUCF-HYU (Industry-University Cooperation Foundation Hanyang University)
Resumen de: US20260245041A1
Disclosed is a graph neural network (GNN) and transformer-based task planner. A task planning method performed by a task planning system may include constructing a task planner model based on a graph neural network and a transformer; and generating a task plan from a goal instruction of a task and history information of a scene graph through the constructed task planner model.