Absstract of: EP4814896A2
0001 A computer-implemented method of generating multimodal data. The method comprises using a token generation neural network to generate an output sequence of multimodal tokens, and in response to a next multimodal token being a start-of-image token, generating an image by performing a reverse diffusion process conditioned on features representing the current output sequence of multimodal tokens obtained from the token generation neural network and appending a sequence of image tokens representing the image to the current output sequence of multimodal tokens as subsequent multimodal tokens in the output sequence of multimodal tokens.
Absstract of: EP4814972A1
Some embodiments are directed to generating molecular structures, which may include progressively denoising a spatial density function by iteratively applying a trained neural network. The neural network is trained to transform noisy spatial density functions into structured spatial density functions based on training data comprising molecular structures represented as spatial density functions. After denoising, a sum of predefined density distributions is fitted to the denoising result to obtain the molecular structure.
Absstract of: EP4815476A2
0001 A system and method of re-identifying players in a broadcast video feed are provided herein. A computing system retrieves a broadcast video feed for a sporting event. The broadcast video feed includes a plurality of video frames. The computing system generates a plurality of tracks based on the plurality of video frames. Each track includes a plurality of image patches associated with at least one player. Each image patch of the plurality of image patches is a subset of the corresponding frame of the plurality of video frames. For each track, the computing system generates a gallery of image patches. A jersey number of each player is visible in each image patch of the gallery. The computing system matches, via a convolutional autoencoder, tracks across galleries. The computing system measures, via a neural network, a similarity score for each matched track and associates two tracks based on the measured similarity.
Absstract of: EP4814895A1
The disclosure relates generally to methods and systems for cross-domain based change detection of region due to an event. Conventional techniques on cross-domain change detection (CDCD) rely on transformation-based approaches involving two tasks: image translation from one modal to another modal (SAR to optical) and then performing the change detection (CD) in one of the translated modals (for example, in SAR or optical). The methods and system of the present disclosure propose a deep-learning (DL) architecture called ReFUjetNet leverages generic embeddings from the extensively pre-trained domain adaptive foundation model alongside locally trained embeddings. The ReFUjetNet efficiently fuses these embeddings and employs a Kolmogorov-Arnold Network (KAN)-based convolutional neural network (CNN) classifier to generate a binary change matrix, which is then joined with another binary change matrix derived from the Gabor jet-based dissimilarity checker, resulting in the final binary change map.
Absstract of: NZ773408A
A milk analyser (400) comprising a milk analysis unit (402) having an analysis modality wherein the milk analysis unit (402) further comprises a milk classification system (404) having an imaging device (4042, 4044) configured to image milk for generation of digital image data; a processor (3044) of a computing device (304) which is adapted to execute a program code to implement a deep learning neural network classifier trained using labelled milk images from milk within the classes into which the imaged milk may be classified and operable to generate a classification of the imaged milk; and a controller (3066) configured to output a control signal in dependence of the generated classification to control a sample intake (4022) to regulate the supply of milk to the analysis unit (402).
Absstract of: US20260289307A1
0000 Techniques for training a neural network having a plurality of computational layers with associated weights and activations for computational layers in fixed-point formats include determining an optimal fractional length for weights and activations for the computational layers; training a learned clipping-level with fixed-point quantization using a PACT process for the computational layers; and quantizing on effective weights that fuses a weight of a convolution layer with a weight and running variance from a batch normalization layer. A fractional length for weights of the computational layers is determined from current values of weights using the determined optimal fractional length for the weights of the computational layers. A fixed-point activation between adjacent computational layers is related using PACT quantization of the clipping-level and an activation fractional length from a node in a following computational layer. The resulting fixed-point weights and activation values are stored as a compressed representation of the neural network.
Absstract of: US20260292179A1
0000 A neural network-based image processing method and apparatus according to an embodiment of the present invention may: acquire a feature tensor from an input image by using a first neural network including a plurality of neural network layers; acquire a symbol tensor by performing quantization on the acquired feature tensor; and generate a bitstream by performing entropy encoding on the basis of the symbol tensor.
Absstract of: US20260289809A1
0000 An electronic device mounted on a fixed or a movable apparatus is provided. The electronic device may comprise a neural processing unit (NPU), including a plurality of processing elements (PEs), configured to process an operation of an artificial neural network model trained to detect or track at least one object and output an inference result based on at least one image acquired from at least one camera; and a signal generator generating a signal applicable to the at least one camera.
Absstract of: US20260289299A1
0000 A method for training a neural network to control a technical system. The method includes ascertaining a partitioning of a control function operating on an input space into a set of first affine functions, each operating on a particular first subset of the input space; training the neural network to approximate the control function in a plurality of iterations, comprising, for each iteration: ascertaining, for each first subset, how the neural network partitions the first subset into second subsets such that it behaves like a particular second affine function on each second subset; ascertaining, for each of the second subsets, an approximation error between the neural network and the control function; ascertaining an overall approximation error between the neural network and the control function from the ascertained approximation errors; and adapting the neural network to reduce the overall approximation error.
Absstract of: US20260290075A1
A machine learning model (MLM) may be trained and evaluated. Attribute-based performance metrics may be analyzed to identify attributes for which the MLM is performing below a threshold when each are present in a sample. A generative neural network (GNN) may be used to generate samples including compositions of the attributes, and the samples may be used to augment the data used to train the MLM. This may be repeated until one or more criteria are satisfied. In various examples, a temporal sequence of data items, such as frames of a video, may be generated which may form samples of the data set. Sets of attribute values may be determined based on one or more temporal scenarios to be represented in the data set, and one or more GNNs may be used to generate the sequence to depict information corresponding to the attribute values.
Absstract of: US20260290329A1
0000 A low power analog Long Short-Term Memory (LSTM) recurrent neural network has an input layer, an array of Adaptive Filter Unit for Analog LSTM, a linear projection layer, and an output layer. The output layer has multiple nonlinear amplifiers, a nonlinear element with a sigmoidal input-output characteristic function, and a time-constant adjustable, nonlinear, low pass filter that provides the memory function of the LSTM. The LSTM memory is used with mismatch-robust weights determined by learning by computation of optimal weights values, wherein the objective function minimizes misdetection probability, and used to process a signal to detect events.
Absstract of: US20260288904A1
Apparatuses, systems, and techniques to train a machine-learned model. In at least one embodiment, a plurality of training clients each obtain an exclusive right to update a model in turn, and each client trains said model with training data not accessible to other training clients.
Absstract of: US20260289238A1
0000 Aspects of the present disclosure involve a system comprising a computer-readable storage medium storing a program and method for video synthesis. The program and method provide for accessing a primary generative adversarial network (GAN) comprising a pre-trained image generator, a motion generator comprising a plurality of neural networks, and a video discriminator; generating an updated GAN based on the primary GAN, by performing operations comprising identifying input data of the updated GAN, the input data comprising an initial latent code and a motion domain dataset, training the motion generator based on the input data, and adjusting weights of the plurality of neural networks of the primary GAN based on an output of the video discriminator; and generating a synthesized video based on the primary GAN and the input data.
Absstract of: US20260289296A1
Apparatuses, systems, and techniques are presented to generate image or video content. In at least one embodiment, one or more neural networks are used to generate one or more time-lapsed images of a second object based, at least in part, on one or more images of a first object.
Absstract of: US20260290539A1
An estimate of a functional capacity such as VO2Max is made by applying the vital signs of a monitored human to a trained encoding neural network producing a cardio profile vector. The vector is applied to a trained functional capacity (VO2Max) neural network to estimate the functional capacity. Once estimated, an action is taken.
Absstract of: US20260289736A1
0000 Apparatuses, systems, and techniques to blend two or more images based on confidence values of objects within said two or more images. In at least one embodiment, one or more confidence values in one or more images are generated using one or more neural networks that are used, for example, to blend two or more images to be displayed.
Absstract of: US20260289724A1
0000 Apparatuses, systems, and techniques for texture synthesis from small input textures in images using convolutional neural networks. In at least one embodiment, one or more convolutional layers are used in conjunction with one or more transposed convolution operations to generate a large textured output image from a small input textured image while preserving global features and texture, according to various novel techniques described herein.
Absstract of: US20260285315A1
0000 In various examples, a three-dimensional (3D) intersection structure may be predicted using a deep neural network (DNN) based on processing two-dimensional (2D) input data. To train the DNN to accurately predict 3D intersection structures from 2D inputs, the DNN may be trained using a first loss function that compares 3D outputs of the DNN—after conversion to 2D space—to 2D ground truth data and a second loss function that analyzes the 3D predictions of the DNN in view of one or more geometric constraints—e.g., geometric knowledge of intersections may be used to penalize predictions of the DNN that do not align with known intersection and/or road structure geometries. As such, live perception of an autonomous or semi-autonomous vehicle may be used by the DNN to detect 3D locations of intersection structures from 2D inputs.
Absstract of: US20260289563A1
Cryptographic control regulates access by DevOps computing to an API service through an API gateway. A processor applies a long short term memory neural network to a plurality of performance metrics of the application programming interface service in real-time, the plurality of performance metrics including a time to first hello world value, a request per minute count, an average latency value, a maximum latency value, an errors per minute count, an application programming interface uptime value, a memory usage value, and a central processing unit usage value. The long short term memory neural network produces a forecast output predicting a future value of at least one of the plurality of performance metrics. When the forecast output indicates the future value is outside a predetermined performance range, the processor disables access in accordance with a smart contract bound to a non-fungible token. A non-fungible token repository data structure is dynamically updated.
Absstract of: US20260289311A1
A method and apparatus with object estimation model training is provided. The method include generating a cross-correlation loss based on a first feature vector, generated using an interim first neural network (NN) model provided an input based on first input data about a target object, and a second feature vector generated using a trained second neural network provided another input based on second input data about the target object; and generating a trained first NN model, including training the interim first NN model based on the cross-correlation loss.
Absstract of: US20260289053A1
The invention relates to a shale fracture seismic identification method based on a 3D U-Net convolutional neural network combined with ant tracking. The method establishes a geological model of shale fractures using single-well data and performs seismic forward modeling to determine the advantageous frequency band for fracture identification. Spectral-peak decomposition is applied to obtain the advantageous frequency-band data volume, which is processed using a 3D U-Net convolutional neural network and ant-tracking computation to generate a 3D U-Net Ant Tracking volume. The results are verified using microseismic data, and along-layer attributes of the 3D U-Net Ant Tracking volume are extracted to determine the regional planar distribution characteristics of shale fractures. The method effectively reduces exploration costs by combining single-well and seismic data, significantly improves seismic resolution through integrated application of seismic forward modeling, spectral-peak decomposition, and advantageous frequency-band data computation.
Absstract of: US20260289995A1
A computer implemented method of performing single pass optical character recognition (OCR) including a fully convolutional neural network (FCN) engine including at least one processor and at least one memory, the at least one memory including instructions that, when executed by the at least processor, cause the FCN engine to perform a plurality of steps. The method includes preprocessing an input image having machine printed characters; extracting two or more image features using a plurality of convolutional layers in the FCN engine; transforming feature maps from the plurality of convolutional layers to a size of the input image using a transposed convolution layer of the FCN engine; generating a multi-channel pixel-level probability output based on the transformed feature maps; scaling the multi-channel pixel-level probability output; generating word boxes for the machine printed characters; determining each character of the machine printed characters; and transmitting, for display, the machine printed characters.
Absstract of: US20260289366A1
0000 Example aspects of the present disclosure provide systems and methods to learn machine-learned model parameters for models of quantum computing systems. In particular, example aspects of the present disclosure are directed to systems and methods to learn a deep neural network configured to predict parameter values for a physical model that models quantum dynamics of interactions between one or more qubits of a quantum gate and one or more two-level-system (TLS) defects during operation of the quantum gate through use of an evolutionary algorithm.
Absstract of: US20260290366A1
0000 In a speech enhancement method performed by an electronic device using artificial intelligence according to an embodiment of the present disclosure, the method may include: determining a hidden representation for a sequence of an input speech signal in a first neural network; determining speech recognition information based on the hidden representation; receiving the speech recognition information and an input data block in a second neural network; and generating an output data block by combining the speech recognition information and the input data block. (FIG. 8)
Nº publicación: US20260289948A1 24/09/2026
Applicant:
SMITH & NEPHEW INC [US]
SMITH & NEPHEW ORTHOPAEDICS AG [CH]
SMITH & NEPHEW ASIA PACIFIC PTE LTD [SG]
Smith & Nephew, Inc.
Smith & Nephew Orthopaedics AG
Smith & Nephew Asia Pacific Pte. Limited
Absstract of: US20260289948A1
Some examples are directed to a processor-implemented methods of training a neural network for soft-tissue labeling of computed tomography (CT) images, neural networks trained according to one or more of the processor-implemented methods, processor-implemented methods of soft-tissue-labeling a CT image, and systems for labeling soft-tissue in computed tomography images.